Skip to content

Production readiness checklist

Use this checklist as an approval record, not as a list one operator silently ticks. The service owner defines business-critical paths, the platform owner proves infrastructure recovery, the security owner approves identity and exposure, and application owners prove data and rollback. Assign every section before the readiness review.

A production-ready installation has known owners, tested recovery, observable customer paths, and evidence that common failures do not require improvising on the host. Passing a build or seeing green cards in the UI is necessary but not sufficient.

  • Canonical HTTPS URL, cookies, WebSockets, and OAuth redirects work.
  • PostgreSQL, Redis, Relay identity, encryption keys, and uploaded artifacts are persistent and backed up.
  • Relay owns public 9443/tcp and recovers after restart.
  • License state and required entitlements are healthy.
  • Two independent administrator recovery paths exist.
  • MFA and group scopes are tested with real non-admin users.
  • API/OAuth/MCP credentials are least-privilege and inventoried.
  • Every role reports expected capabilities and update channel.
  • Offline behavior, daemon restart, node restart, and reconnect are tested.
  • Firewalls allow only documented traffic.
  • Storage Nodes reach ghcr.io, and an LXC Storage Node has a loop-device pool sized for every managed instance and concurrent backup.
  • Docker Nodes that hold Availability leases run the lease watchdog and reach every Relay.
  • Deploy, health check, rollback, logs, backup, and restore are proven.
  • Managed database bindings survive Opfield, Relay, daemon, node, and workload recreation tests.
  • No owner credentials are present in workload configuration.
  • DNS, TLS renewal, ingress migration, maintenance mode, and upstream failure are tested.
  • Notifications and status-page communication reach their intended recipients.
  • Update and rollback runbooks have owners.
  • Disk, certificate, database, queue, Relay, and node alerts are active.
  • An incident drill has been completed without relying on undocumented shell changes.

Do not approve production from visual inspection alone. Preserve evidence for:

  • installation version and image digests;
  • backup completion and a recent restore test;
  • external DNS/TLS/HTTP verification;
  • non-admin permission tests;
  • Node restart and reconnect tests for every deployed role;
  • Opfield, Relay, daemon, and workload restart behavior;
  • database binding query results before and after recreation;
  • alert open/recovery and notification delivery;
  • update and rollback rehearsal;
  • unresolved exceptions with owner and expiry.

Deployment is no-go when administrator recovery is untested, persistent keys are not backed up, Relay has a single unknown recovery path, node capability errors are ignored, customer traffic cannot be verified externally, or a database binding requires manual container/role repair.

An accepted temporary exception must state the customer impact, compensating control, owner, deadline, and rollback trigger. “Works on the current host” is not a production-readiness argument.

After enabling business traffic, watch Route health, TLS, Relay pressure, node freshness, workload restarts, database connections, disk growth, notification queues, and audit activity. Keep the rollback owner available through the observation window and avoid unrelated platform changes.

Hold the review against a named release and installation. For each area, record pass, accepted exception, or no-go, plus an evidence link and owner. Do not carry evidence forward from an older release when the relevant lifecycle, daemon, Relay, database, or ingress component changed.

Use one row for every acceptance check so a less experienced reviewer can see what was tested and a power user can reproduce it:

Area Owner Check performed Result Evidence Exception or follow-up
Example: public application Application team External DNS, TLS, and primary user path Pass Task, request, or report link None

Do not write only “works” or “checked.” The evidence should identify the exact resource, version, test location, time, and expected result.

Start with the three journeys that would create immediate business impact: administrator recovery, public customer traffic, and application access to persistent data. Then review supporting capabilities such as builds, Pages, notifications, AI, and optional integrations according to what the installation will actually use. Untested unused features do not block launch; enabled critical features do.

  1. A new authorized user signs in, completes MFA, and is denied an out-of-scope resource.
  2. A supported workload is deployed, observed, restarted, and rolled back without host edits.
  3. A public Route is verified from an external network with the expected certificate and health behavior.
  4. A managed Node disconnects and reconnects without losing desired state.
  5. Opfield restarts while customer workloads remain in their documented state.
  6. A backup is restored into a clean compatible environment and required secrets decrypt.
  7. Alerts open and recover during a controlled failure, and the intended recipient receives them.

The approval should state the scope that was tested. Do not claim that the whole product is validated when only one Node role, database engine, runtime profile, or update path was exercised.