Node updates and offline behavior
Node updates and outages are continuity events, not merely daemon restarts. Previously applied services may continue locally while Opfield loses the ability to change or observe them in real time. The operating goal is therefore to preserve the durable Node identity, restore the authenticated session, and prove that every dependent resource reconciled—not only that the status returned to green.
The service owner decides the maintenance window and acceptable interruption. The platform owner confirms version compatibility, host access, recovery evidence, and post-update acceptance. Success means a fresh connection, inventory, metrics, capabilities, and representative workload paths, with no unexplained interrupted operation.
Operator details: update and reconnect
Section titled “Operator details: update and reconnect”Check compatibility before updating Opfield, Relay, and daemons. Apply signed daemon updates through the Node lifecycle and verify reconnect, version, inventory, logs, health, and managed bindings.
When a node goes offline, workloads may continue locally but Opfield mutations are disabled. Diagnose host availability, daemon service, system time, certificate identity, DNS, and outbound Relay access. Avoid deleting the Node record while the old daemon may reconnect.
After recovery, verify desired-state reconciliation for Routes, workloads, database bindings, certificate material, and monitoring. A green connection alone does not prove every dependent resource recovered.
Crash-safe update launcher
Section titled “Crash-safe update launcher”In 2.10, Docker, nginx, monitoring, and the Relay supervisor gained a separate launcher. When launcher-managed, an update preserves the previous binary and an on-disk journal. The candidate must report local readiness and pass a 30-second stability window before commitment; failed candidates can roll back. The service manager supervises the launcher, which supervises the daemon process.
Local readiness does not prove Relay connectivity or application health. Continue to verify version, reconnect, Node capabilities, and a real operation after updating. If launcher bootstrap is unavailable and the daemon runs directly, launcher rollback protection is unavailable.
In 2.11, a candidate daemon that declares control-connection readiness must also receive its first command from Opfield within 3 minutes, otherwise the launcher rolls the update back. Daemon updates of Docker, nginx, and Relay supervisor Nodes also replace the launcher; in 2.10 it stayed at the version it was installed with. The updated daemon stages a new launcher once any running launcher trial has finished, the launcher tries it at its next start, and keeps it once the trial is stable, so the installed launcher follows the daemon one update behind.
Certificate renewal
Section titled “Certificate renewal”Opfield renews its internal certificates while it runs, without a restart, and records each renewal in the audit log:
| Certificate | Validity | Renewal |
|---|---|---|
| Opfield listeners and the local Relay’s service and client certificates | 365 days | Checked hourly and renewed 30 days before expiry; existing connections keep the old certificate until they reconnect. While unused enrollment tokens exist, renewal waits until the last 7 days, because installer commands pin the Opfield certificate fingerprint; commands generated before a renewal must then be generated again |
A certificate supplied through GRPC_TLS_CERT |
Set by you | Never renewed by Opfield; a daily warning is logged |
| Remote Relay server certificates | 365 days | Checked hourly while the Relay supervisor is connected and renewed 60 days before expiry, at most once a day |
| Daemon client certificates on every Node | 365 days | Checked hourly by the daemon and renewed once a third of the lifetime remains, about 120 days before expiry. A failed attempt is retried after 5, 10, 20, and 40 minutes, then hourly. The daemon reconnects to present the new certificate |
| Registry proxy certificate on Docker Nodes | 2 years | Checked when the daemon starts and every 12 hours, and reissued from the same CA once a third of the lifetime or 30 days remain; served without a restart |
| Managed storage and managed database TLS certificates | Set by Opfield | Checked hourly and reloaded in place without a restart; see managed storage and managed databases |
A Node must stay connected to renew its certificate. Because renewal starts months before expiry, a Node certificate with less than 30 days left means renewal is failing: Opfield raises a warning alert at 30 days and a critical alert at 7 days. Make sure the Node can reach Opfield and runs a current daemon; after its certificate expires, a Node must be enrolled again.
The system CAs that issue these certificates are valid for 10 years and have no automatic rollover in 2.11. Certificates they issue end with the CA, and Opfield alerts 730, 365, 180, 60, 30, and 7 days before a system CA expires. See Internal PKI.
Enrollment tokens
Section titled “Enrollment tokens”Enrollment tokens are single-use and expire after 7 days. For a Node that is still pending, New enrollment token issues a replacement installer command. An enrolled Node cannot be re-enrolled in place; only Relay members support re-enrollment. See Relay and Relay Pool.
Daemon updates
Section titled “Daemon updates”Update one Node from its detail page with Update, or several at once with Nodes > Update Nodes. The dialog lists every Node whose daemon is older than the latest release of its type and updates the selected ones together. Relay Nodes are not listed: they update with the Relay Pool (Settings > General > Update Relay Pool), which drains them one at a time. Opfield refuses a daemon update for a Node that is not connected, for a Relay Node, and for a Node that already runs that release or a newer one.
- Update state. A Node keeps its update state in the list and on its page, without an update badge or an offline flicker while it updates. A failed, rolled-back, or timed-out update keeps its reason on the Node, shown under Update Available, and a Node that comes back on its old version is reported as such right away.
- Lease peers. Nodes that share a lease-mode Availability policy restart one after another; a queued update shows Queued with what it waits for. See Updates and mixed versions.
- Opfield restarts. After an Opfield restart during Node updates, running updates wait for their Nodes to reconnect and finish or fail on time, and queued updates are taken up again. Opfield checks for daemon updates when it starts and after an Opfield update.
- Broken connections. A daemon that staged a self-update restarts into it even when its connection to Opfield broke during the download.
- Installers. Node and Relay setup commands come from the running Opfield installation’s own release and are checked against that release’s checksums, and daemon installers pin the newest daemon of that installation’s release line, so an Opfield installation never hands out installers from a newer development line.
Before an update
Section titled “Before an update”- Read the component release notes and compatibility floor.
- Confirm the installation update channel is intentional: Stable for production releases or Preview for explicitly accepted prereleases.
- Verify the Node is online and not running a conflicting lifecycle operation.
- Record its current version, capabilities, active workloads, and recent errors.
- For a Relay Pool, use the Relay Pool update, which drains, updates, and verifies one member at a time.
- Ensure an operator can reach the host if automatic recovery fails.
Opfield and daemons verify signed update manifests and checksums. Do not replace a failed signed update with an unsigned binary copied from another host.
Update verification
Section titled “Update verification”After the daemon restarts, verify more than its version:
- authenticated reconnect and fresh last-seen time;
- complete capability report;
- inventory and metrics refresh;
- role-specific service health;
- pending task reconciliation;
- Routes, certificates, workloads, Compose projects, and database bindings owned by that Node;
- alerts generated or cleared as expected.
Offline behavior by responsibility
Section titled “Offline behavior by responsibility”An offline control connection does not automatically stop host services. nginx, containers, managed databases, and other previously applied services can continue from local state. Opfield prevents mutations that require a live owner and shows cached inventory only for diagnosis.
Database bindings can continue through a healthy Relay and database path during an application-only Opfield restart; Relays keep admitting new connections on their last signed policy for the policy lease. A Relay outage is different: it interrupts new private-link admissions and managed-node control traffic even when local workloads continue. When Opfield returns, daemons reconnect within 15 seconds.
Docker daemon restarts
Section titled “Docker daemon restarts”Restarting or updating the Docker daemon does not stop containers, and it interrupts traffic to them only briefly:
- The daemon connects to Opfield first and registers its Relay endpoints again, within about a second, once Opfield’s fresh grants arrive. Secure Runtime (gVisor) is verified again in the background afterwards. Until that finishes, the Node reports the last result verified for the same
runscversion, or an unknown status with the reasonverification_pendingon the first start of a new version, so creating secure workloads does not flap. - A graceful restart or update is announced. Before it stops, the daemon tells the Relays that its endpoints are restarting, finishes the Secure Link requests in flight (at most 1 second), and closes connections that sit idle between requests. The Relays keep its registrations for up to 15 seconds until the next process registers, and Ingress Nodes hold new connections to it for up to 8 seconds instead of failing them, so a restart shows as a delay of a few seconds rather than errors. An Availability replica with another serving replica does not wait: its requests go to the other replica at once. This needs Relays and nginx daemons of this release; with older Relays, the Ingress Node still holds new connections for up to 3 seconds.
- The holder of an Availability lease keeps serving across a restart on the same boot: its lease is still valid, so its copy takes traffic again as soon as the new process registers.
- A daemon that crashes, or whose host becomes unreachable, is not held: the Relays drop its connection within about 3.5 seconds, and requests move to other placements where there are any.
- After a host reboot, each Secure Link is restored on its own: a link whose target did not come back stays unbound until its target runs, and does not block the others.
- A hung Docker Engine (dockerd) does not stop traffic to containers and Deployments that are running: Secure Link, managed database, and storage connections wait at most 300 milliseconds for it and then use the last address it named, links stay bound when it does not answer a sync, and Deployment routers use neither Docker’s DNS nor an access log. An application that writes every request to its container output can still stall while dockerd is frozen.
- A Storage or database Node starts even when one of its managed containers is missing; that storage or database shows as not running instead of taking the Node offline.
- Warnings of every daemon component, such as refused link connections, reach the Node’s logs in Opfield, not only the host journal.
Nodes that vote in or can hold data-plane failover leases are updated one at a time per Availability policy: a daemon update of such a Node may wait for its lease peers, and the Node shows as updating meanwhile. See Updates and mixed versions. On Docker Nodes that can hold leases, the lease watchdog runs as its own service, gateway-lease-watchdog, and updates itself.
A Node rolled back to any daemon release since 2.10.0 still starts: daemons keep the state files an older release reads next to their own. Daemons replace their state files durably (write, flush, rename), so a kill -9 or a power loss never leaves a partial file behind.
Docker daemon 2.11.1 and the shared connector
Section titled “Docker daemon 2.11.1 and the shared connector”The 2.11.1 Docker daemon serves every database binding, storage link, and container link of the Node through one shared secure-link connector, with no per-link containers and no listeners on the host; see Shared connector and link networks.
- Database bindings move to the connector after the update: Opfield recreates each bound workload once, one at a time on each Node, the next once the previous runs healthy. A Deployment rolls out blue/green without downtime; a standalone Container or a Compose service restarts once.
- Storage links created before 2.11.1 switch to the connector after the update, and Opfield recreates each linked workload once right after its link switches. Links created on 2.11.1 need no recreate.
- Rolling back the Docker daemon to 2.11.0 switches the links back automatically, again with one recreate per workload. At most four workloads per Node are recreated at once, the daemon’s limit on concurrent commands, so on a Node with many linked Deployments the later ones reconnect a little later.
- Connector updates keep open link connections for up to 30 minutes, then close them once; clients reconnect to the new connector.
- Container links need the 2.11.1 Docker daemon on both Nodes and show update required until then.
- Link network range: new link networks take a /26 each from
docker.secure_links.subnet_poolin/etc/docker-daemon/config.yaml,10.213.x.x(a /16) by default. The daemon skips subnets that Docker networks or the host’s routes already use; change the range when it overlaps a network the Node reaches through its default route, then restart the daemon.
nginx daemon restarts
Section titled “nginx daemon restarts”Restarting or updating the nginx daemon on an Ingress Node does not close its Secure Link sockets. The launcher that supervises the daemon keeps a copy of every listening socket and hands them to the next daemon process; with systemd it also keeps them in the unit’s file descriptor store, so they survive systemctl restart. While no daemon process runs, new connections wait in the socket’s queue and are served by the next process. On shutdown the daemon stops accepting, finishes requests in flight for up to 1.5 seconds, and closes connections idle between requests so that nginx reconnects into the queue.
- The update from 2.10 is visible once. The old daemon hands nothing over when it stops, so that one update interrupts Secure Link routes through the Node for about a second. From then on, restarts and updates keep the sockets, including the restart in which the launcher tries its new version: while an older launcher that keeps no sockets still runs, the daemon stores its listeners in systemd’s file descriptor store itself.
- Requests that run longer than the 1.5-second drain, and long-lived streams such as WebSocket or server-sent events, are cut at a restart.
- On Nodes without systemd (OpenRC), restarts for updates and after a crash are covered by the launcher, but a service restart is not. Routes whose configuration still uses a loopback TCP listener instead of a Secure Link socket refuse connections during a restart, as before.
Secure Link sockets are created under a temporary name with their final owner and mode and then renamed, so nginx never finds a missing socket or one it may not use; before a configuration is loaded, every socket it references already listens.
When Opfield starts before an Ingress Node has reconnected, changes for that Node, such as rewritten routes after a template update, are left to the Node’s reconnect sync instead of failing.
Recovery sequence
Section titled “Recovery sequence”- Restore host power and network reachability.
- Verify system time, DNS, and Relay endpoint connectivity.
- Inspect the daemon systemd unit and bounded logs.
- Confirm the daemon certificate and Node identity were not replaced.
- Wait for reconnect and fresh inventory.
- Review failed or interrupted operations before retrying them.
- Verify each dependent resource family owned by the Node.
Do not delete an offline Node merely to clear a warning. Deletion changes durable ownership and can leave the still-running old daemon unable to reconcile safely.
Interrupted operations
Section titled “Interrupted operations”An offline transition can leave an operation whose intent was accepted but whose final host result was not acknowledged. After reconnect, Opfield reconciles supported interrupted state from the daemon and resource inventory. Review the operation before starting another action: repeating a create, update, migration, or deletion blindly can conflict with work that already completed on the host.
Compare the operation record with the actual owner resource. A Container may be running even if the last UI progress was interrupted; a Compose Project may have applied a revision without returning its final acknowledgement; an update may have restarted the daemon before the control session returned. Use the refreshed snapshot and role-specific state as evidence, then retry only the supported continuation or recovery action.
Force-cancelling a Task stops or abandons control-plane work where supported; it does not guarantee that every host-side subprocess or already-applied runtime change was reversed. Verify the resource after cancellation and perform explicit cleanup when the operation reports it is required.
Replacement versus recovery
Section titled “Replacement versus recovery”Recover the existing identity when the host and its persistent daemon material still exist. Replace the Node only when the host is intentionally rebuilt or its identity cannot be recovered. Before replacement, inventory Routes, workloads, database instances, build assignments, certificates, volumes, and Relay assignments owned by the old Node and move or retire them through their supported lifecycle.
Never run the same daemon identity on an old and replacement host simultaneously. If the original host might return, disable or uninstall it before enrolling the replacement. A duplicate identity can produce conflicting inventory and operation acknowledgements even when the two hosts have different addresses.
Planned offline test
Section titled “Planned offline test”Before relying on a Node in production, perform a bounded recovery exercise: restart the daemon, confirm local services behave as documented, wait for reconnect, and verify fresh inventory plus one representative dependent resource. For Relay and database paths, test new connections as well as existing ones. Record the observed recovery time and the manual access path needed if automatic reconnect fails.