Auto-Deploy Pipeline
Auto-Deploy Pipeline Linear: OPE-275 Problem Deploying to production is manual and error-prone: SSH in, git pull, rebuild all containers (causing downtime), ...
Auto-Deploy Pipeline
Linear: OPE-275
Problem
Deploying to production is manual and error-prone: SSH in, git pull, rebuild all containers (causing downtime), manually update Caddy config. Steps are easy to forget — especially Caddyfile updates, which have been a recurring bug source.
The admin-sidecar (backend/admin_sidecar/main.py) currently handles update
orchestration inside Docker Compose, but it has a fundamental flaw: it can’t
reliably restart itself (the process dies mid-execution when its container is
recreated), and if the Docker stack crashes, the orchestrator dies with it.
Architecture: Host-Level Deploy Agent + GitHub Webhook
Update orchestration moves to a host-level systemd service that lives outside
Docker Compose. GitHub sends a webhook on push to main; the deploy agent
receives it, verifies the signature, and orchestrates the full update. A polling
fallback catches missed webhooks.
The admin-sidecar is stripped down to read-only endpoints (logs, version, health).
GitHub (push to main)
│
▼ POST /deploy/webhook/github
┌────────────────────────────────────┐
│ Caddy (host, systemd) │
│ Routes /deploy/* → port 9000 │
│ Routes everything else → 8000 │
└────────────────────────────────────┘
│ │
▼ ▼
┌──────────────┐ ┌─────────────────────────────────┐
│ Deploy Agent │ │ Docker Compose Stack │
│ (systemd, │ │ │
│ port 9000) │──▶│ api, app-*, workers, cache, │
│ │ │ vault, cms, admin-sidecar, ... │
│ - webhook rx │ │ │
│ - git pull │ └─────────────────────────────────┘
│ - compose │
│ build/up │ ┌─────────────────────────────────┐
│ - caddy │──▶│ /etc/caddy/Caddyfile │
│ reload │ │ (validate + reload) │
│ - health chk │ └─────────────────────────────────┘
└──────────────┘
Component Split
| Component | Runs as | Responsibility |
|---|---|---|
Deploy Agent (deployment/deploy-agent.py) |
systemd service on host (port 9000) | Webhook receiver, update orchestration, Caddy reload, health checks, polling fallback |
Admin Sidecar (backend/admin_sidecar/main.py) |
Docker container (port 8001) | Read-only: /admin/logs, /admin/version, /health |
CLI (openmates server auto-update) |
One-time setup command | Install systemd service, generate webhook secret, configure GitHub |
Why Host-Level (Not Container)
- Survives stack crashes — can bring Docker Compose back up
- Can restart everything including the admin-sidecar itself
- Directly manages Caddy (also on host) — no volume mount gymnastics
- No Docker socket inside containers — smaller attack surface
- systemd
Restart=alwayshandles agent crashes automatically
Deploy Flow
- PR merged to
main→ GitHub sends webhookPOSTtohttps://api.openmates.org/deploy/webhook/github - Caddy routes the request to the deploy agent on port 9000 (bypasses Docker stack entirely)
- Deploy agent verifies HMAC-SHA256 signature (
X-Hub-Signature-256header) using shared secret - Agent checks
refmatchesrefs/heads/main(configurable branch) - Update sequence runs in background thread:
a.
git pull(120s timeout) b.docker compose buildall services (600s timeout) — no downtime, old containers keep running c. Optional: stop cache container +docker volume rmcache volume (clean restart) d.docker compose up -d— swap all containers to new images (~30s downtime) e.docker compose up -d vault-setup cms-setup— re-init secrets and schema f. DetectSERVER_ENVIRONMENT, copy correct Caddyfile to/etc/caddy/Caddyfile, validate, reload g. Health check: pollhttp://localhost:8000/healthevery 5s for up to 60s - Returns 202 Accepted immediately; caller polls
/deploy/statusfor progress
Polling Fallback
When GITHUB_POLL_INTERVAL_MINUTES > 0, a background thread periodically:
- Calls
GET https://api.github.com/repos/{owner}/{repo}/commits/{branch}(unauthenticated, 60 req/hour limit) - Compares returned SHA with local
git rev-parse HEAD - Triggers update if they differ and no update is in progress
Recommended: 10-minute interval (6 requests/hour, well within rate limits).
Deploy Agent Endpoints
| Endpoint | Method | Auth | Description |
|---|---|---|---|
/webhook/github |
POST | HMAC-SHA256 (GitHub signature) | Receive GitHub push webhook, trigger deploy |
/status |
GET | X-Admin-Log-Key header |
Last update time, status, step log |
/health |
GET | None | Liveness probe for systemd |
Configuration
All config via environment variables, loaded from .env by systemd EnvironmentFile:
GITHUB_WEBHOOK_SECRET — HMAC shared secret (generated by CLI setup)
GITHUB_DEPLOY_BRANCH — branch to deploy (default: main)
GITHUB_POLL_INTERVAL_MINUTES — polling interval, 0=disabled (default: 0, recommended: 10)
GITHUB_REPO — owner/repo for polling (e.g. glowingkitty/OpenMates)
DEPLOY_AGENT_PORT — listen port (default: 9000)
GIT_WORK_DIR — project root
SERVER_ENVIRONMENT — production/development (for Caddyfile selection)
COMPOSE_FILE — path to docker-compose.yml (relative to GIT_WORK_DIR)
CLEAR_CACHE_ON_UPDATE — true/false (wipe Dragonfly cache volume)
CACHE_VOLUME_NAME — e.g. openmates-cache-data
ADMIN_LOG_API_KEY — for /status endpoint auth
Systemd Service
File: deployment/openmates-deploy-agent.service
Follows the same pattern as scripts/agent-trigger-watcher.service:
Type=simple— script runs in foregroundRestart=always,RestartSec=10— auto-restart on crashEnvironmentFile— loads.envfor configUser=<USER>,Group=docker— Docker access without rootAfter=network.target docker.service— starts after Docker is ready
Caddy Routes
Added to both deployment/prod_server/Caddyfile and deployment/dev_server/Caddyfile,
before the catch-all abort:
@deploy_webhook {
path /deploy/webhook/github
method POST
}
handle @deploy_webhook {
reverse_proxy localhost:9000
}
@deploy_status {
path /deploy/status
method GET
}
handle @deploy_status {
reverse_proxy localhost:9000
}
These routes bypass the Docker stack entirely — even if every container is down, the webhook endpoint remains reachable.
CLI Setup
openmates server auto-update setup:
- Generates 32-byte hex webhook secret
- Appends config to
.env:GITHUB_WEBHOOK_SECRET,GITHUB_DEPLOY_BRANCH,GITHUB_POLL_INTERVAL_MINUTES,GITHUB_REPO,DEPLOY_AGENT_PORT - Installs systemd service: copies unit file, fills
<USER>placeholder, runssystemctl daemon-reload && systemctl enable --now openmates-deploy-agent - Prints GitHub webhook setup instructions (payload URL, secret, content type, events)
- Optionally creates webhook via
gh apiif GitHub CLI is available
openmates server auto-update status:
- Calls
GET http://localhost:9000/status - Shows last update time, status, deploy branch, poll interval
Admin Sidecar Changes
The admin-sidecar loses its update endpoints but keeps read-only functions:
| Endpoint | Status |
|---|---|
POST /admin/update |
Removed — moved to deploy agent |
GET /admin/update/status |
Removed — moved to deploy agent |
GET /admin/logs |
Kept — Docker log access |
GET /admin/version |
Kept — commit SHA, tag, branch |
GET /health |
Kept — liveness probe |
Docker Compose env vars removed from sidecar: SERVICE_UPDATE_ALL, SERVICE_UPDATE_TARGET,
SERVICE_UPDATE_EXTRAS, CLEAR_CACHE_ON_UPDATE, CACHE_VOLUME_NAME.
Self-Hosted Users
The same deploy agent works for self-hosted installations. The CLI auto-update setup
command provisions everything needed. Future enhancement: GHCR pre-built images so
self-hosted users can pull instead of building locally (admin-sidecar already supports
this via its Docker Hub mode).
Files
| File | Action | Purpose |
|---|---|---|
deployment/deploy-agent.py |
New | Host-level webhook receiver + update orchestrator (stdlib Python, zero deps) |
deployment/openmates-deploy-agent.service |
New | Systemd unit file |
deployment/prod_server/Caddyfile |
Edit | Add webhook + status routes to port 9000 |
deployment/dev_server/Caddyfile |
Edit | Add webhook + status routes to port 9000 |
backend/admin_sidecar/main.py |
Edit | Remove update endpoints, keep logs/version/health |
backend/core/docker-compose.yml |
Edit | Remove update env vars from sidecar block |
frontend/packages/openmates-cli/src/server.ts |
Edit | Add auto-update subcommands |
Night Deploy Window + Post-Deploy Smoke Tests
Incident context (2026-04-04): A manual server restart during active user sessions caused 4 bug reports (OPE-316 through OPE-319) from a single user within 40 minutes. Cache was wiped, sync broke, AI lost conversation context, responses rendered as empty, and the user was charged 1,687 credits for broken responses (balance went to -17).
Requirements
-
Scheduled deploy window — auto-deploy only between 2:00–4:00 AM CET (lowest traffic). PRs merged outside this window are queued until the next window.
-
Pre-deploy active user check — query connected WebSocket count before starting. If > 0 active users, delay deploy by 15 min and re-check (max 3 retries, then deploy anyway with alert).
-
Post-deploy smoke test — after containers are healthy, run a minimal E2E test:
- Login with test account
- Send a chat message and verify AI response is received
- Verify WebSocket connection is established
- If smoke fails → alert (email + push), do NOT auto-rollback (manual investigation)
-
Smoke test runs on production — verify the actual production deployment, not dev.
-
Nightly test suite runs after deploy —
tests.py run --dailyPlaywright suite should run after a successful deploy + smoke pass, not on a fixed schedule. This catches regressions from the deploy. -
Production E2E tests — extend the nightly suite to run key specs against production (login, chat, shared chat, payments). Currently all E2E tests only run against dev.
Deploy Sequence
1. PR merged to main → queued for next deploy window
2. Deploy window opens (2:00 AM CET)
3. Check active WebSocket connections → delay if users online
4. Run deploy: git pull → build → up -d → vault/cms setup → Caddy reload
5. Health check: poll /health for 60s
6. Post-deploy smoke test: login + send message + verify response
7. If smoke passes → run nightly E2E suite (dev + production specs)
8. If smoke fails → alert admin, log failure, do NOT rollback automatically
Additional Config
DEPLOY_WINDOW_START_HOUR — earliest hour for deploy in server timezone (default: 2)
DEPLOY_WINDOW_END_HOUR — latest hour for deploy (default: 4)
DEPLOY_MAX_ACTIVE_USERS — max WebSocket connections before delaying (default: 0)
DEPLOY_SMOKE_TEST_ENABLED — run smoke test after deploy (default: true)
Implementation Order
- Deploy agent script + systemd unit + Caddy routes (core auto-deploy)
- Night deploy window + active user check
- Post-deploy smoke test (login + chat + WS verification)
- Production E2E specs in nightly suite
- Strip update logic from admin-sidecar (clean separation)
- CLI setup command (self-hosted experience)
Out of Scope (Future)
- GHCR image publishing — for self-hosted users who shouldn’t build locally
- Auto-rollback — smoke test provides visibility now; auto-revert to previous git SHA later
- Staging environment
- Database migration automation