Workload Availability (HA)
Availability adds placements on independent Docker nodes to an existing Container, Deployment, or Compose Project. Opfield preserves the resource identity, configuration, and relationships while creating and replacing runtime instances according to its availability policy.
The feature is available on Business and Enterprise. Manage instances through the owning resource rather than as individual Containers: the policy owns placement count, generations, routing membership, and cleanup.
Failover runs in one of two ways. When every Docker Node, Ingress Node, and Relay of the workload runs a release that supports it, the Nodes fail over by themselves, without Opfield; see Data-plane failover. Otherwise Opfield replaces lost placements itself (backend failover), as described on this page.
Choose a mode
Section titled “Choose a mode”| Mode | Behavior | When to use it |
|---|---|---|
| Replicated | Maintains 2–32 serving placements, at most one per node | The application supports concurrent instances |
| Failover | Maintains one serving placement and creates a replacement after node loss | One serving instance with recovery on another node is required |
Without Availability, the resource retains its ordinary single-node lifecycle. Each Compose placement contains the whole project; Opfield does not distribute individual services of one instance across nodes.
Replica count is manual. Metric autoscaling, multiple placements of the same workload on one node, and opportunistic rebalancing of healthy instances are outside this model.
Prerequisites
Section titled “Prerequisites”Enabling Availability requires at least two online compatible Docker nodes. Provide enough capacity for the desired replica count and temporary placements during updates.
The workload must have no configured or observed mounts, including named, external, read-only volumes, or host bind mounts. For Compose, the whole project is checked. Persistent data must live in separate services; Availability does not copy local data between nodes.
Opfield pins images by repository and digest in its internal registry and pre-pulls them on eligible standby nodes. Verify registry access, Relay and Secure Link paths, and dependent managed databases.
Enable Availability
Section titled “Enable Availability”- Open the Container, Deployment, or Compose Project and its Availability section.
- Select Check eligibility and resolve incompatibilities.
- Turn on Enable and choose Mode. For Replicated, set Serving placements.
- Under Eligible nodes, choose All compatible nodes or Selected nodes with an explicit node list.
- Review replacement and rollout settings, then select Save.
- Follow the operation and Placements list until the requested serving count is reached.
- Verify a real request through the Route, database access, and application logs.
Changing the switch alone does not apply the policy: Save is required. A successful save means desired state was accepted, not that instance creation has completed.
Enabling on a running workload
Section titled “Enabling on a running workload”Enabling Availability on a workload that already serves does not interrupt it and never starts a second copy beside it. The running Container, Deployment, or Compose Project becomes the first placement on the Node where it runs:
- A Container is never recreated. Managed database networks it needs are attached while it runs.
- A Deployment or Compose Project that already runs the image Opfield mirrored into its registry is adopted as it runs. An adopted Deployment keeps its router’s port binding, and an adopted Compose Project its published ports, until a rollout recreates them.
- A Deployment or Compose Project that cannot be adopted as it runs, because it needs the managed database connection settings or runs another image, is replaced without a serving gap where the policy allows two copies (Replicated, or Failover with Allow available mode): a Deployment deploys the new configuration to its inactive slot and switches after readiness; a Compose Project first gets a copy on another Node, the Routes move to it, then the original is recreated and, in Failover mode, the extra copy is removed. If no other Node can serve the copy, the original stays as it is and the enable retries.
- A Failover singleton with Allow available mode off never runs two copies, so it takes a short planned gap instead: a Deployment creates its new slot with the image already pulled before the old slot stops, and switches back if the new slot does not become ready; a Compose Project pulls every image before any service is recreated.
The Node where the workload runs keeps serving it even when Priority mode puts another Node first; failback moves it afterwards.
Routes keep reaching the original workload through their own Secure Link until a placement serves through its member link. Only then does Opfield switch the Route to the members; it keeps the old link for another 5 seconds so that requests nginx’s previous workers still hold can finish, and then removes it.
Replacement and rollout settings
Section titled “Replacement and rollout settings”- Replacement grace — time to wait after losing the node control connection before creating a replacement. The default is 15 seconds; image preparation, startup, and health checks take additional time.
- Maximum unavailable — how many placements a planned update may make unavailable simultaneously.
- Maximum surge — how many temporary placements may exist above the desired count. These require additional capacity.
- Drain interval — time between removing a placement from new routing and stopping it, allowing existing connections to finish.
Priority mode
Section titled “Priority mode”By default, Opfield places serving instances on eligible nodes by free capacity. Priority mode makes either mode prefer nodes in a fixed order instead. Turn it on below Eligible nodes and order the list with the up and down buttons: the first node is Primary, the next ones are Backup 1, Backup 2, and so on. The list holds only eligible nodes; nodes that become eligible later are added at the end. Select Save to apply the order.
Serving placements go to the first available nodes in the order. When a higher-priority node returns and stays healthy for Return to the primary after (300 seconds by default, 0–3600), Opfield runs a Failback operation: it starts the workload on that node, switches traffic to it, and then drains and removes the backup placement, so a placement keeps serving throughout. Turning priority mode on or changing the order moves the workload the same way. The Availability summary shows the Primary node and where the workload is Serving from, marked On backup while a backup serves.
- A node that goes offline or reports an error restarts its delay, so a flapping node does not take traffic back.
- Failback runs only while the policy is healthy and no other operation of the same policy is active; replacement, rollout, and scaling of that policy come first. Operations of different policies run in parallel, so a failback never waits for another workload’s heal.
- A failed failback leaves the backup serving and is tried again after 15 minutes. Use Retry in Operations to try sooner.
- Failback needs Maximum surge of at least 1 so that a placement keeps serving during the move; in Replicated mode, Maximum unavailable of at least 1 is also enough. Otherwise Check eligibility warns, and the workload stays on the backup.
- When the nodes run the failover themselves (data-plane failover), a failback is a handoff instead: the serving node stops its copy before the next node starts its own, because the default strict partition mode never runs two copies at once. In Failover mode the workload does not serve while the old copy stops and the new one starts and becomes ready, usually 5–15 seconds; in Replicated mode the other replicas keep serving.
- After an Opfield restart, every node’s healthy time starts again, so no failback runs until the delay has passed.
In the API and MCP (manage_docker_availability), the policy fields are priorityMode, nodePriority (node IDs in order, primary first), and failbackDelaySeconds.
Routes and database bindings
Section titled “Routes and database bindings”Proxy Hosts, Additional Routes, and Advanced Secure Links can target the logical workload. Opfield projects healthy placement endpoints and balances new ingress connections using least connections. Existing connections do not move between replicas.
A placement takes traffic only once its workload is ready: its image health check reports healthy if the image has one; otherwise its port accepts connections. For a Deployment, the application of the active slot must answer through the Deployment router, not only the router itself. Route health is checked through the placements’ Secure Links, so a healthy Availability Route is shown online.
A placement that does not serve, such as a standby or a stopped copy, keeps its link closed on the Ingress Node in every mode, so no request reaches it: nginx fails the connection before it sends anything and moves on to the next placement, which is safe for every method.
If a placement fails anyway, the Ingress Node retries the next placement: on connection errors for every request, and on 502, 503, and 504 responses for requests that are safe to repeat, such as GET and HEAD, up to 5 tries within 10 seconds. POST, PATCH, and LOCK requests are not sent again once they reached a placement. A placement that failed is skipped for one second; every placement is also listed as a backup that is never skipped, so the Route never runs out of upstreams while one placement serves.
When the host of a placement becomes unreachable, the Relays drop its connection within about 3.5 seconds and report it as not ready, and the Ingress Node closes its link. Until then, a new connection to it costs at most 2 seconds before nginx tries the next placement. Requests already in flight to it are retried on another placement when they are safe to repeat; a POST already sent gets 502. A placement’s link gives up at once only when a Relay answered about that placement and another placement serves. When no Relay can be reached at all, for example while the only Relay restarts, or no other placement serves, new connections are held like on any Secure Link: up to 3 seconds, or 8 seconds for an announced restart.
Managed database bindings receive per-placement connections. Keep relationships attached to the logical resource rather than temporary Container names. To diagnose a particular instance, use the placement selector in logs, console, and monitoring where available.
A container link to an Availability workload follows its healthy placements: each consumer Node prefers a placement on its own Node, the link moves away from a placement Availability judges unhealthy and returns once it passes its health checks again. A link to an Availability Deployment must use one of the Deployment’s ports, because the Deployment router forwards only those.
Updates and replica health
Section titled “Updates and replica health”A rolling update that fails for good, or keeps failing for 3 attempts or 15 minutes, is rolled back automatically: replicas that were already updated return to the previous image and settings, and the policy keeps the original error. Changing a setting no longer recreates every replica.
A rollout updates every copy, including the original workload on its first Node, and completes only once each Node runs the new version of its copy. A rollout that an Opfield restart or update interrupts is not failed or rolled back: it resumes when Opfield is back, with the image and settings it was started with.
Availability works with multi-platform images on Docker 29 and later, which use the containerd image store. After an Opfield update, a heal that finds the internal registry still starting retries instead of failing.
Opfield checks replica health every 30 seconds. A replica that is missing, stopped, restart-looping, or unhealthy on two checks in a row is removed from routing and healed; it returns after two healthy checks. A standalone Container that already has a Route can enable Availability, and Availability can be disabled even when no replica is healthy.
Node loss and recovery
Section titled “Node loss and recovery”In lease mode, the Nodes replace a lost holder themselves; see Data-plane failover. On backend failover, a Node whose control connection drops is shown offline after 5 seconds, and its placements leave routing after 10 seconds, so a short Relay restart does not move traffic. After Replacement grace, Opfield starts a prepared standby if the policy has one, and otherwise creates a replacement on an available eligible node. When the lost Node reconnects, Opfield heals the policy immediately.
In both modes, a policy with fewer serving placements than requested and eligible online capacity is healed without an edit: Opfield checks every 15 seconds, outside a Node’s offline grace, and whenever a Node connects. An update that sets the current serving count or node selection again also repairs missing placements. If the requested count is still not restored, inspect the operation phase, node compatibility and capacity, image pull, application health, and private dependencies.
On backend failover, an unreachable host may still run its old process. Opfield excludes stale generations from managed routing and reconciles them on reconnect. This is not physical node fencing or a single-writer guarantee for an external system; only data-plane failover makes a cut-off Node stop its own copy.
Backend failover requires a functioning Opfield control plane. Availability does not provide HA for Opfield itself, nginx, database engines, registry storage, or shared volumes. Verify Relay-path resilience and each dependency separately.
After recovery, verify serving count, real customer traffic, database access, and stale-placement cleanup. Do not manually delete child Containers to force the operation to complete.
Stop and start
Section titled “Stop and start”Stop on a workload with Availability stops the copies that are running and keeps every placement: standbys stay prepared, and the policy stays healthy while the workload is stopped. Start starts the copies that served last, never a standby, and keeps the standbys, so no placement is removed, recreated, or failed by a stop and a start. The original workload is never deleted: not when a lease closes, not by stale-placement cleanup, and not as a surplus standby. On a Node that is no longer selected it stays, stopped. A Docker daemon that is busy with long-running commands delays a stop, start, or heal (AVAILABILITY_DAEMON_BUSY, retried automatically) instead of marking placements failed. A start, stop, or restart requested while the policy’s previous operation still runs is queued and runs after it; Containers, Deployments, and Compose Projects behave the same.
A Stop sent right after Availability was enabled, before the Nodes took over failover, keeps every copy stopped: no Node starts a copy of a stopped workload by itself.
Operations of several policies
Section titled “Operations of several policies”Each policy runs one operation at a time, in order. Operations of different policies run in parallel: heal, scale, rollout, enable, and cleanup operations run for at most 4 policies at once, while failback, start, stop, restart, and disable never wait for that limit and run beside them.
Disable Availability
Section titled “Disable Availability”Turn off Enable and save the change. In Disable Availability, choose Surviving placement, enter the workload name for confirmation, and select Disable and keep one.
The surviving placement keeps running and becomes the ordinary workload as it runs; disabling does not restart or recreate it:
- The original Container, Deployment, or Compose Project is adopted as it runs, and gets its own restart policy back. A Compose Project whose containers were prepared by Availability shows as drifted on the Compose page until your next apply.
- A Deployment copy on another Node is adopted as it runs as well.
- A Container or Compose Project copy on another Node runs under a name Availability generated. It is replaced by a copy with the workload’s original name on that Node, which starts before the old copy is removed, so traffic has no gap.
If the policy is in lease mode, the disable first closes the lease and keeps the surviving copy running through the close. While the disable runs, no heal, scale, failback, or start of the policy runs (AVAILABILITY_DISABLE_IN_PROGRESS), and a disable that has to wait keeps the status disabling. A retry keeps the surviving placement chosen first.
Wait for the operation and the removal of the other placements to finish. Verify the surviving instance, Routes, and database bindings: the resource should continue its ordinary single-node lifecycle.