Administration
For the team that runs an Oqim deployment: the platform admin console, the oqimctl command, configuration, backups, health checks, metrics and logs.
Admin console#
Platform administrators manage the whole deployment at /admin, or on an admin. host routed to it. Opening it takes a platform-admin user, a console session (API keys can't use it) and two-factor authentication: an administrator without it is asked to enroll on the first visit, and every new session asks for a code.
- Every admin action is written to the audit log with the actor type
PLATFORM_ADMIN, including in the affected organization's trail. - Suspending an organization pauses its running and scheduled campaigns, stops its messaging and makes it read-only; its members see your reason in their console. Reactivating doesn't resume those campaigns; the organization does.
- A campaign you pause, or an account you disable, can't be resumed by its organization. The organization is notified and told to contact support. Restore the account from the console; resume the campaign with
POST /api/v1/admin/campaigns/{id}/resume. - Disabling a proxy needs a reason, which the organization sees. Oqim stops connecting through the proxy at once: the account using it stops sending and its campaigns pause, and it never falls back to a direct connection. The organization can give the account another proxy or a direct connection; only you can enable the proxy again.
- New organizations need an owner who already has an Oqim sign-in.
- Settings changes reach every service within about five seconds.
Automatic brakes#
Every 30 seconds the scheduler looks for signs that messages are reaching people who didn't ask for them. The campaign rules pause the running campaign, and every rule opens an abuse report for you to review. The campaign thresholds are under Platform settings → Abuse protection.
The organization is notified of a pause and can resume the campaign. After an automatic pause, only what happens since then counts toward the same rule, so a resumed campaign is judged on its new sends: opt-outs from messages sent before the pause don't pause it again.
oqimctl#
oqimctl performs operator tasks and reads the same environment variables as the services. With Docker Compose, run it in the worker container, which has SESSION_KEY: docker compose exec worker oqimctl ….
After create-admin, sign in to the console, open /admin and set up two-factor authentication.
Configuration#
Services read their configuration from environment variables and report every missing required one together at startup.
Also available: SESSION_TTL_HOURS (console session lifetime, default 336), LOG_LEVEL=debug, and WORKER_ID to name a worker process. The web app takes API_URL at build time: where it forwards /api/v1.
Hosts#
One web deployment serves three surfaces:
Sign-in, invitation and opt-out pages work on every host. Locally, /admin and /docs work directly. To share one sign-in between the console and admin hosts, set COOKIE_DOMAIN to the parent domain and list both hosts in ALLOWED_ORIGINS.
infrastructure/nginx/oqim.conf is an example edge configuration: it terminates TLS for all three hosts and forwards everything to the web app, which proxies /api/v1 to the API. Whatever you put in front, allow request bodies of at least 25 MB for imports and attachments, and turn response buffering off for /api/v1/events/stream.
Backups and recovery#
On a server deployed with deploy.sh (see deploy/README.md), the database is dumped with pg_dump in its custom format into .deploy/backups/ next to the checkout:
./deploy.sh backupmakes a dump now, namedoqim-YYYYMMDD-HHMMSS-manual.dump.- Every deploy makes one before it applies migrations, named
oqim-YYYYMMDD-HHMMSS-before-BUILD.dump. The last 10 dumps are kept (KEEP_BACKUPSchanges that), and./deploy.sh --no-backupskips the dump. - The dumps stay on the server, so copy them somewhere else on a schedule. They cover Postgres only, not attachment files: on the local volume, back up
oqim_attachmentsyourself; on S3, turn on the bucket's versioning or your provider's backups.
Restoring a dump
Stop the services that write to the database, restore with pg_restore, then start them again:
Rolling back
./deploy.sh rollback runs the previous build again (or ./deploy.sh rollback TAG a given kept one) without rebuilding. It doesn't undo migrations: the older build runs against the newer schema, which is why migrations only ever add. To get the data back as it was before a deploy, restore the dump that deploy made.
Health checks#
Both are served by the API at its root, outside /api/v1. Workers report health through their heartbeat, visible under Workers in the admin console: online within 20 seconds of the last heartbeat, degraded within 90 seconds or while erroring, offline after that.
Metrics#
The API, workers and scheduler expose Prometheus metrics on METRICS_ADDR: :9101, :9102 and :9103 by default. docker compose --profile monitoring up starts a Prometheus and a Grafana that scrape them.
Alerts worth having
Logs#
Every service writes structured JSON to standard output, one event per line, tagged with service and env. The API logs each request with its method, route pattern, status, duration, request ID and client IP, plus the actor and organization when authenticated. Set LOG_LEVEL=debug for more detail.
Errors the API hides from clients (internal_error) are logged in full with the request path, so search the logs by path and time when a user reports one. See Security for what Oqim stores.