Production deployment
Deploy immutable, tested artifacts from a reviewed commit. Do not build from a mutable working tree on the server and do not make schema rollback depend on an old binary understanding a new incompatible migration.
Prerequisites
- all repository working trees are clean and the release commit or tag is recorded;
- external credentials required by the release have been rotated or verified;
- a restorable PostgreSQL backup and restore test are current;
- migration preflight checks pass against an upgrade fixture;
- API, worker, both SPAs, SDK, docs and website checks pass with pinned tools;
- the API and SDK contract versions are compatible;
- Stripe test-mode webhook and reconciliation exercises pass for billing changes.
The docs and website use Node 22.17.0 and npm 11.7.0. Use npm ci, which installs the exact lockfile, rather than resolving dependencies during a release.
The production Caddy binary must include the dns.providers.cloudflare module used for wildcard portal certificates. A stock Caddy binary cannot load api/deploy/Caddyfile. Verify the exact binary and service files on a Linux host before reload:
cd api
CADDY=/usr/local/bin/caddy
"$CADDY" list-modules |
awk '$0 == "dns.providers.cloudflare" { found=1 } END { exit !found }'
set -a
. /etc/microauth/caddy.env
set +a
"$CADDY" validate --config deploy/Caddyfile --adapter caddyfile
systemd-analyze verify deploy/microauth-api.service \
deploy/microauth-worker.service deploy/caddy.serviceLocal and CI gates
cd docs
npm ci
npm run check
cd ../website
npm ci
npm run checkThe docs check validates the committed OpenAPI JSON, bundles the pinned Scalar reference, builds VitePress and verifies internal links and assets. The website check builds Astro and runs the same static link and asset validation.
API and SDK pipelines additionally run their unit, integration, race, lint, type and package checks. The API deployment workflow checks out the exact SHA whose CI run succeeded and builds that SHA. Manual dispatch records and builds the selected SHA. Automatic production deployment accepts only a successful push run for main from the API repository itself; pull-request and fork runs are never eligible.
Fresh-database verification
Run migration and usage integration tests against a disposable PostgreSQL server. MA_TEST_DATABASE_URL is an administrative connection to a database server, not the application database. Each test creates a uniquely named database, migrates it from empty, checks behavior, and drops it.
cd api
export MA_TEST_DATABASE_URL='postgres://postgres:postgres@127.0.0.1:5432/postgres?sslmode=disable'
go test -race -count=1 ./internal/database \
-run '^TestMigrationsFromScratchAndCoreInvariants$'
go test -race -count=1 ./internal/modules/sdkapi \
-run '^TestApplyUsageIntegration$'The migration test applies the complete migration set twice, verifies that the schema is clean at the expected version, and exercises core ownership constraints. The SDK integration test creates a second fresh database and checks request counting, billable charging, item idempotency, cap rejection, external-credit replay safety, and 200 concurrent accounting attempts. A skipped test is not a pass; CI sets the same MA_TEST_DATABASE_URL name for both jobs.
Usage accounting guarantees
Each accepted SDK usage item commits its idempotency receipt, exact monthly tenant/customer counters, hourly usage bucket, prepaid balance debit, and credit-ledger entry in one PostgreSQL transaction. Ambiguous client timeouts are retried with the same idempotency key; a committed retry returns duplicate without charging twice. Rejected transactions leave no receipt or counter mutation.
The FastAPI SDK batches usage reports (requests merge into counted items of up to 500, flushed at least every 5 seconds), so portal usage figures can trail live traffic by a few seconds while ingestion load stays a tiny fraction of request volume: a 500-request burst is one API call, one item, one receipt and one accounting transaction.
Usage charges accumulate into one credit_ledger row per customer per UTC day (migration 0020). Billing history growth is bounded by time, not request volume, and the billing page shows daily totals (labeled with the UTC traffic day) instead of a row per charge. The daily cleanup job compacts historical per-item usage rows into daily rows (amounts preserved) and purges hourly usage_records past the 366-day analytics horizon, so no billing table grows without bound.
Usage reports commit in chunked transactions (100 items per transaction, one savepoint per item), so a 500-item report costs a handful of commits instead of 500 sequential transactions while keeping exact per-item accepted/duplicate/rejected semantics.
Deploy note: migration 0019's backfill briefly blocks usage ingestion while it aggregates usage_records (SDK items return retry and are redelivered); apply it off-peak on large installations.
Monthly counters are authoritative O(1) rows, not asynchronous estimates. This keeps quota and platform-cap checks exact without repeatedly scanning the hourly usage table. Concurrent requests are admitted through a per-tenant gate before taking a database connection, preventing lock waiters from exhausting the pool. Snapshot requests use a short cache and single-flight build so cold starts cannot open one repeatable-read transaction per caller. Policy-write races between concurrent snapshot/verify transactions (serialization failures and duplicate-policy conflicts) are retried with fresh transactions instead of surfacing as 500s.
Database pool isolation
The API process opens two PostgreSQL pools from MA_DATABASE_URL:
- Interactive pool (
MA_DB_POOL_INTERACTIVE, default 8): portal host resolution, portal APIs, and SaaS management traffic. - SDK pool (
MA_DB_POOL_SDK, default 12): all/sdk/*ingestion and snapshot traffic, with a 15-second server-sidestatement_timeout.
An SDK ingestion burst from a hot tenant can therefore only exhaust the SDK pool; portal host lookups and dashboards keep their own connections. Size both values so their sum stays within the database's connection budget across all API replicas (the worker uses 10 more).
Migration policy
Use expand-and-contract migrations:
- add nullable columns, tables and compatible indexes;
- deploy code that writes both old and new shapes where needed;
- backfill and verify;
- switch reads after every running version understands the new shape;
- remove old fields in a later release.
Run migrations once under an explicit lock. A failed migration stops the release before binaries are switched. Never claim a binary rollback can undo an incompatible schema change.
The production scripts run microauth-api -migrate from the staged release before activation. The application migrator serializes migration work and returns nonzero on a dirty or failed migration. Do not use a down migration as an automatic deploy trap. Restore a backup only as an explicit incident decision, after stopping all writers and accounting for data created since the backup.
Deployment sequence
- Put the worker in a safe drain state when its job schema changes.
- Apply compatible database migrations.
- Deploy and start the worker version that understands new inbox and reconciliation jobs.
- Deploy the API, then wait for
GET /readyzto prove PostgreSQL and Redis readiness.GET /healthzonly proves the process can answer HTTP. - Run contract smoke tests against the API.
- Deploy the tenant portal and SaaS SPA versions compatible with that API.
- Publish the SDK only after the API accepts its protocol.
- Regenerate and check OpenAPI locally, then deploy docs and the marketing site.
Use atomic artifact replacement, root-owned read-only binaries, checksums and an unprivileged deploy identity with narrowly scoped service permissions. Validate service and reverse-proxy configuration before a controlled reload.
The API workflow and api/deploy/deploy.sh follow this activation sequence:
- upload artifacts and
SHA256SUMSto a unique release directory; - verify checksums and run migrations from the staged API binary;
- copy the currently active API and worker into that release's
previousdirectory; - install both new binaries under hidden
.nextnames, then rename them into place before restarting either service; - restart API and worker once, verify both systemd units, and poll
/readyz; - on any post-swap failure, stage and rename both saved binaries back, restart both units, and leave the failed release directory available for diagnosis. If a complete previous pair does not exist, stop both services instead of claiming a partial rollback.
Each rename is atomic on the host filesystem. The pair is not one filesystem transaction, so both .next files must be fully installed before the first rename and neither service is restarted between renames. The manual script activates the portal only after API readiness by replacing one symlink to its immutable release directory.
Smoke tests
/readyzis successful and the worker can claim a harmless test job;- login, CSRF, workspace roles and portal team scoping behave correctly;
- two browser tabs can use different portal workspaces;
- snapshot, verify and one idempotent usage report work;
- retrying that usage item returns
duplicateand does not change balance; - runtime pricing loads, and its failure state offers retry;
- a signed Stripe test event reaches the durable inbox and is processed;
- docs API reference loads the committed local snapshot;
- canonical, favicon, Open Graph, sitemap and internal links resolve.
Rollback
Stop the rollout when readiness or smoke tests fail. Restore the previous application artifacts only when the new migration remained backward compatible. Otherwise roll forward with a corrective build. Financial events remain in the durable inbox during rollback and are replayed only through idempotent handlers after compatibility is restored.
Before a production release, verify that previous contains both active binaries and record their checksums. After an automatic rollback, confirm both services and /readyz; the trap deliberately does not claim the database was rolled back. If compatibility is uncertain, keep traffic stopped and roll forward with code that understands the migrated schema.
Record the deployed commit, migration versions, artifact checksums, smoke-test result and operator. Do not include secret values in deployment logs.