Incident runbook
This runbook helps the incident lead protect customer service, establish which subsystem owns the failure, and choose a reversible recovery path. The incident lead owns priorities and communication; the platform operator owns Opfield and managed-node evidence; application and database owners validate their services. Do not let the person typing commands become the only person deciding risk.
Success means customer impact is understood, further change is controlled, the authoritative runtime state is reconciled, and recovery is verified from the user-facing path. A green Dashboard alone is not recovery evidence.
First five minutes
Section titled “First five minutes”- Record start time, affected customer path, and recent changes.
- Check Opfield, PostgreSQL, Redis, and Relay health.
- Identify whether impact is control plane, ingress, one node, one workload, or a shared dependency.
- Freeze unrelated changes.
- Use maintenance mode or a status-page incident when customer impact is confirmed.
Relay unavailable
Section titled “Relay unavailable”Inspect the Relay service, identity volume, database reachability, image version, and public 9443/tcp. In Settings > Relay, check each member’s policy-trust and certificate state: Opfield recovers local Relay trust automatically, and a remote member marked for re-enrollment or with an expired certificate can be recovered with Re-enroll or Renew certificate; see Relay recovery in 2.11. Do not move the port back to the application or bypass authorization with a new public listener.
Node offline
Section titled “Node offline”Existing workloads may continue. Availability policies in lease mode move the Node’s slots to other candidates by themselves; check Lease mode and the holders in the workload’s Availability summary before acting. Check host power/network, daemon service, time, certificate identity, and outbound Relay reachability. Do not mutate from stale inventory.
Bad release
Section titled “Bad release”For Deployments, roll back to the previous healthy slot. For Pages, move the Tag. For Git-built containers, redeploy the last approved digest. If the target reports BUILD_ROLLOUT_IN_PROGRESS, wait for the running rollout to finish or fail before redeploying. Preserve logs and operation history.
Database binding failure
Section titled “Database binding failure”Check Relay, both nodes, database health, binding desired state, the binding’s link on the target Node’s secure-link connector, and the engine principal. Retry reconciliation before manual identity changes; there is no per-binding connector container or host listener to restart or replace.
Use the layer-by-layer checks in Application database bindings and preserve the failed Task before retrying. When clients report Connection terminated unexpectedly or ECONNRESET under load, especially during a rollout, read the link runtime first: growing admission rejects with link_limit mean the containers’ pools exceed the link’s 64 connections; see Link capacity and connection pools.
Closeout
Section titled “Closeout”Confirm recovery from an external client, clear maintenance intentionally, document the root cause and detection gap, and create follow-up work for every temporary mitigation.
Control-plane unavailable
Section titled “Control-plane unavailable”Check the Opfield container/process, PostgreSQL, Redis, persistent volumes, disk pressure, migrations, and application logs. While the API still answers, the assistant or an MCP client with diagnostics:view can read all of this with manage_gateway_diagnostics, and diagnostics:logs adds the container logs; see Opfield diagnostics. Preserve the existing encryption keys and database before attempting a replacement instance. If customer workloads continue, avoid broad host changes while restoring the control plane.
First determine whether Relay is still reachable. If Relay is healthy, already applied nginx configuration, containers, databases, and permitted existing private streams can continue while new management and authorization decisions wait. Relays keep admitting new Secure Link and database-link connections on their last signed policy for the policy lease, Availability policies in lease mode keep failing over, and the public status page is served from the Ingress Node’s cache. If the only Relay is local to the failed Opfield host, Secure Links and managed-node sessions lose their transport as well; ordinary host-local workloads may still run, but existing Relay-dependent streams are not guaranteed. Independent external Relay members reduce this shared failure only when Nodes can actually reach and use them.
After Opfield returns, verify sign-in, decryption, Relay health, background jobs, update discovery, notifications, and fresh node snapshots. Do not assume daemons reconciled merely because the Dashboard loads.
Use Availability, compatibility, and limits to classify the affected paths and Updates and backups before restoring a replacement instance or changing persistent state.
Ingress outage
Section titled “Ingress outage”Follow Ingress troubleshooting from DNS through network, TLS, nginx apply, Route health, and upstream. Use maintenance mode when the virtual host should remain stable during repair. Preserve generated-config validation errors and external request evidence.
Workload operation stuck
Section titled “Workload operation stuck”Inspect the durable Task, owning Node connection, daemon operation history, runtime resource, and request/operation ID. A browser timeout does not prove the command failed. Reconcile the first operation before starting a duplicate create, recreate, migration, or delete.
If Force Cancel is available, understand whether it cancels only Opfield tracking or also dispatches cancellation to the owner. After cancellation, refresh runtime state and clear any interrupted-operation marker only through the supported reconciliation path.
Use Tasks, events, and audit to distinguish accepted, running, cancelled, and completed work, then continue with the owning Docker resource.
Database unavailable
Section titled “Database unavailable”Separate engine health, storage, direct publication, managed binding, and application-query failures. Verify the Storage Node, engine container, storage image/mount, monitoring snapshot, Relay path, target Docker Node, the Node’s secure-link connector, binding identity, and application configuration.
Do not delete the database or engine principal as a diagnostic step. Preserve storage and operation history, then retry the narrow failed reconciliation.
Continue with Database operations for engine, storage, backup, and restore checks, or Application database bindings when only the private application path is affected.
Communication and evidence
Section titled “Communication and evidence”Maintain one incident timeline with UTC timestamps, affected resources, user-visible impact, changes, task IDs, request IDs, decisions, and verification. Public updates describe impact and progress without exposing topology or security details.
Recovery gate
Section titled “Recovery gate”Recovery requires:
- the customer-facing path works from outside the managed network;
- Opfield desired state matches current owner state;
- alerts recover for the right reason;
- queued operations and deliveries are understood;
- temporary exposure or bypasses are removed;
- a follow-up owner and deadline exist for every remaining risk.
Decision framework for the incident lead
Section titled “Decision framework for the incident lead”Prefer actions that preserve evidence and reduce blast radius. Pause unrelated automation before restarting shared services. If the control plane is unavailable but workloads still serve traffic, restore management without recreating healthy resources. If ingress is unhealthy, keep the hostname stable with maintenance mode or a tested rollback rather than changing DNS repeatedly. If data integrity is uncertain, stop writes before optimizing availability.
Escalate immediately when the incident involves possible credential disclosure, damaged storage, unknown schema migration state, loss of Relay identity, inconsistent database binding ownership, or a release artifact that cannot be verified. These conditions can turn a routine restart into permanent loss or unauthorized access.
Operator evidence table
Section titled “Operator evidence table”| Evidence | Why it matters |
|---|---|
| External DNS, TLS, and application response | Confirms actual customer impact |
| Opfield, PostgreSQL, Redis, and Relay health | Separates control-plane and data-plane failure |
| Node last seen, capability state, and daemon logs | Identifies the runtime owner and stale inventory |
| Task, operation, and request IDs | Prevents duplicate or ambiguous mutations |
| Last approved version and artifact digest | Provides a known rollback target |
| Database and volume backup status | Defines safe recovery options |
After recovery, perform a short review while evidence is still available. Record the initiating change, why detection did or did not work, which manual steps were needed, and which verification would have caught the problem earlier. Convert temporary instructions into a tested runbook change rather than preserving undocumented shell history.