Data-plane failover
Workload Availability can fail over in two ways. On backend failover, Opfield notices that a Node is gone and creates a replacement; this needs a running Opfield. In data-plane failover (lease mode), the Docker Nodes, Ingress Nodes, and Relays of the workload decide among themselves who serves it, so a dead or cut-off Node is replaced within about 45 seconds even while Opfield is stopped, being updated, or unreachable.
Lease mode is part of Business and Enterprise. It needs no switch: a policy uses it as soon as every participant supports it and has run steadily for 2 minutes. Participants are the Docker Nodes that can run the workload (candidates), the Ingress Nodes of its Routes, and the Relays that carry its traffic. They must advertise the availability_lease_v2 capability, which the Docker daemon, nginx daemon, and Relay of Opfield 2.11 do. Components of early 2.11 release candidates that advertise availability_lease_v1 count as outdated.
The Availability summary of the workload shows Lease mode: Lease when the Nodes fail over by themselves, Legacy on backend failover, and Bootstrapping or Closing while the policy switches. Hover the badge for the reason.
How it works
Section titled “How it works”Each serving placement is a slot: Failover mode has one, Replicated mode one per replica. For every slot, the Nodes and Relays of the policy keep a lease, and only the Node that holds it (the holder) runs its copy of the workload.
- Renewal. The holder renews its lease every few seconds with a majority of the policy’s voters. When it cannot, it stops its own copy before the lease can pass to anyone else. The next candidate then takes the free slot and starts its copy, about 45 seconds after the failure.
- Standbys. Opfield prepares up to two standby placements on other candidate Nodes ahead of time: the image is pulled and the container is created but not started, so a takeover only starts it.
- Traffic. Requests reach only the holder of each slot. Every placement’s Secure Link is registered on every Relay that supports lease mode; the links of standbys stay registered but closed until their Node holds the slot, so the new holder is reachable through every Relay right after a takeover. A holder takes traffic only once its workload is ready; see Routes and database bindings.
- Release. A holder that stopped its copy for a local reason (its lease timer, a lost watchdog, a detected freeze, a closed lease) releases the slot as soon as the container is confirmed stopped, so the next candidate starts within seconds instead of waiting for the lease to expire.
- One slot per Node. In Replicated mode, a Node never holds two slots of the same policy. A Node that comes back takes its own previous slot back first.
- Unreachable holders. When a holder’s host stops answering, the Relays drop its connection within about 3.5 seconds, and Ingress Nodes stop sending it requests; see Routes and database bindings. In Replicated mode the other replicas serve from then on; the slot itself moves once the lease has passed to the next candidate.
A Docker Engine (dockerd) that hangs on the holder does not stop traffic to the running copy, and it is not taken as a reason to stop it. If the holder’s lease runs out meanwhile, the lease watchdog stops the copy.
Lease watchdog
Section titled “Lease watchdog”Every Docker Node that can hold a slot runs the lease watchdog, a small service named gateway-lease-watchdog with its own release line. It stops a lease-mode container once its lease deadline has passed, even when the Docker daemon or dockerd hangs. The Docker node installer installs it, on existing Nodes the Docker daemon installs it itself when it runs as root, and the watchdog keeps itself updated. A daemon that cannot install it makes the policy list that Node with watchdog_missing: re-run the node installer there.
A watchdog that is slow for a few seconds (up to about 8) never stops a workload. When its heartbeat has been missing for about 10 seconds, the holder stops its copy and releases the slot to the next candidate. If a Docker daemon is rolled back to a release without lease support, the watchdog removes the deadlines left behind after 10 minutes, so it does not stop containers the older daemon runs.
Voters and witness
Section titled “Voters and witness”Each policy has its own voters: its candidate Nodes, one per physical host, in takeover order, plus witnesses so that the count is odd and at least 3, at most 7. Only Nodes and Relays that advertise availability_lease_v2 and have reported a lease identity vote. A policy with three candidates needs no witness.
A witness is a Relay or Docker Node on a different host from every candidate. Choose it under Witness in the policy: Auto (farthest by latency) picks the eligible member whose shortest round trip to any candidate is the longest, which is least likely to share a site with them; you can also select a specific Relay or Node. The automatic witness is never Opfield’s local Relay while another Relay or Node can witness, because the local Relay stops together with Opfield. It stays until it can no longer vote, so round-trip jitter never changes the voters. The summary warns when the witness is probably on the same site as a candidate (under 2 ms), when no eligible witness exists, or when the configured witness cannot vote and an automatic one is used instead.
Voters must reach each other without Opfield. Every Node of a lease-mode policy connects to every Relay for its lease traffic, and one live Relay is enough while a majority of the voters can reach each other through it. The summary warns “Voter reachability margin is insufficient” when losing one more voter would leave no majority; Opfield’s local Relay does not count toward this margin. A policy with two candidates and a Relay on another host has a margin of one voter.
Network splits
Section titled “Network splits”Allow available mode (partitionMode in the API) decides what happens when the network splits the voters:
- Off (
strict, the default). Two copies of a slot never run at once. Only the side that keeps a majority of voters serves; a holder on the other side stops its copy. - On (
available). Both sides of a split keep serving, and two copies may run for a short time. Never allow it for singletons such as indexers, queue consumers, or scheduled jobs.
Frozen hosts and clock changes
Section titled “Frozen hosts and clock changes”Lease timers use each host’s monotonic clock. A change of the wall clock, such as an NTP correction in either direction and of any size, never stops a copy.
If the holder’s host or virtual machine is paused, frozen, or suspended for more than about 2 seconds, its lease may have passed to another Node meanwhile. After it resumes, the Node sees the time it lost from the clocks of the first Relay or voter it hears from and stops its copy at once, without a grace period when its lease time has run out. The Docker daemon logs this as a fence with the reason host_frozen. A lease the Node acquires after the freeze is never stopped because of it, and shorter pauses stay within the normal timing margins. A holder that reaches nobody after resuming stops its copy on its own lease timer.
Excluded Nodes
Section titled “Excluded Nodes”A condition of one Node never changes the policy’s mode. Instead, the Node is listed under Excluded nodes in the summary (lease.excludedNodes in the API) with one of these reasons:
| Reason | Summary label | Meaning and fix |
|---|---|---|
offline |
offline | The Node has no control connection to Opfield. Restore the host or its daemon. |
watchdog_missing |
lease watchdog not running | The lease watchdog does not run. Start gateway-lease-watchdog or re-run the node installer. |
daemon_outdated |
daemon outdated | The Docker daemon does not advertise availability_lease_v2. Update it. |
identity_pending |
no lease identity yet | The daemon has not reported its lease identity yet, usually right after an update. |
An excluded Node gets no new standby, no failback or other planned move goes to it, and its daemon does not take a slot; the next candidate does. Opfield never cuts off a Node that holds a slot for this reason: a Node without a watchdog stops its own copy, and an outdated holder keeps its slot until it is updated or hands over. An outdated or unidentified Node leaves the voters and the candidate list only after the condition has lasted 2 minutes, so a daemon restart or a rolling update changes nothing. A failback to a Node that was excluded waits Return to the primary after from the end of the exclusion.
Entering lease mode
Section titled “Entering lease mode”A policy on backend failover enters lease mode only when every participant it needs has run availability_lease_v2 for 2 minutes without a break and without a restart: every candidate Docker Node with its lease watchdog and identity, the Ingress Nodes of its Routes, the Relays that carry it, and its witnesses. Until then its reason is participants_settling and names the Nodes or Relays still settling, so a fleet in the middle of an update never switches halfway. After an Opfield restart the 2 minutes start again.
Lease mode also waits for a running enable or rollout of the policy to finish. The first holder of each slot is the Node where the workload serves at that moment, whatever the priority order says; a priority failback moves it afterwards as a planned handoff. If that Node is not ready to take its slot, for example because it finds no copy of the workload, its lease watchdog does not run, or the copy’s restart policy is not no, its Docker daemon logs the reason once per reason.
Entering lease mode does not interrupt the workload. The serving copies keep running and keep their Secure Links: only a lease formed after the switch counts for a slot, so an old lease from an earlier lease period never makes a Node stop its own copy, and the links of serving copies change their role in place instead of being registered again.
Leaving lease mode
Section titled “Leaving lease mode”A policy leaves lease mode only when lease mode has become impossible:
| Reason | Cause |
|---|---|
insufficient_voters |
Fewer members that can vote than a majority needs; update the candidates or add a Relay as witness |
ingress_not_capable |
An Ingress Node of the workload’s Routes runs an nginx daemon without availability_lease_v2 |
relays_not_capable |
A Relay that carries the workload runs a version without availability_lease_v2 |
candidates_not_capable, watchdog_missing |
No candidate can hold a slot, and no slot is held |
controller_unsupported |
This Opfield edition runs Availability failover from the backend only |
signing_key_pending |
No Relay policy signing key can sign lease data yet |
controller_unsupported means a Community build without the lease controller. The license never closes lease mode: no license state change, including expiry, reaches this decision.
Except for disabling Availability and lifecycle operations such as Stop, Start, and Restart, which leave at once, the condition must last 2 minutes without a break. Meanwhile the lease keeps working, the API shows since when in lease.reason.since, and the summary warns: “Data-plane failover is impossible right now: … The policy goes back to backend failover at … unless this is fixed before; the serving copy keeps running.”
Leaving lease mode never stops the serving copy. The policy passes through Closing: Opfield publishes a closed lease that names, for each slot, the Node holding it. Once a majority of the voters confirm the close to that Node, it keeps its copy running without a lease, its watchdog no longer guards that copy, and backend failover adopts it as running, usually within seconds; nothing restarts it and its Secure Links keep serving. While this happens, the summary shows the Node under Kept serving with “confirming” and then “keeps running” (lease.retainedHolders in the API). A holder that cannot reach a majority during the close, because it is cut off, stops its copy like any lease holder, and backend failover starts that slot again once its lease has expired, about 47 seconds after a majority of voters saw the close.
A close keeps every placement, including the original workload: standbys stay prepared, and nothing is removed or queued for cleanup. If a copy has to be started afterwards on backend failover, a prepared standby is started before any new placement is created. Only Stop stops the workload, after the close. Disabling Availability keeps the surviving placement running and removes the other placements itself. When the cause is fixed, the policy returns to lease mode by itself after the usual 2 minutes of settling, and takes the kept standbys as they are.
What Opfield still does
Section titled “What Opfield still does”Opfield plans every move and keeps the records; the Nodes only decide who serves when something fails.
- Operations run in parallel across policies. Each policy runs one operation at a time. Heal, scale, rollout, and similar background operations run for at most 4 policies at once; failback, start, stop, restart, and disable never wait for that limit. A Node that returns can therefore fail back several workloads without queueing behind other policies’ heals.
- Planned moves are handoffs. A failback, drain, manual move, change of the priority order, or rollout moves a slot as a handoff: the serving Node stops its copy and releases the lease, and only then does the next Node start its own. The members’ Secure Links switch roles in place, the old holder’s to standby and the new holder’s to serving, so no link or Relay endpoint is created during a handoff. The slot serves nothing for the old copy’s stop plus the new copy’s start and readiness, usually 5–15 seconds. In Failover mode with Allow available mode off, requests fail for that time; in Replicated mode the other replicas keep serving.
- Healing. A policy with fewer serving copies or standbys than it needs, and eligible online capacity, gets maintenance every 15 seconds and whenever a Node connects. A Node that comes back gets its free slot back: its placement is kept, and a missing one is created again. A Node that ends up holding two slots hands one to a standby elsewhere, with a short gap for that copy.
- Standbys stay as they are. Preparing a standby again, which healing, a reconnect, or maintenance does regularly, changes nothing while its settings, its image (compared by image ID), and the managed database connection settings it would receive are unchanged. A standby whose database connection settings changed, for example after a binding credential was rotated, counts as not ready and is prepared again. A Docker daemon that is busy with long-running commands delays healing (
AVAILABILITY_DAEMON_BUSY, retried) instead of failing placements. - After a partition or outage. When a Node returns, Opfield reads the state of its placements from its daemon: a copy that was stopped becomes a ready standby, and a priority failback runs as after any outage.
- When Opfield returns. A serving copy keeps its Secure Link path while Opfield reconnects its Nodes: if refreshing a member link fails because a Node has not reconnected yet, the link keeps its Relay endpoint and Opfield retries. A Relay that reconnects receives every policy change at once.
- Records. Within seconds of starting, Opfield learns the actual holders from its local Relay, before the Nodes reconnect. An autonomous takeover appears in the audit log as
docker.availability.lease_failover, a planned move asdocker.availability.lease_handoff. Both are dated with the time the new holder acquired the slot, which can be earlier than the moment Opfield noticed it. A slot that lapsed and was taken again by the same Node, because its copy was stopped or cut off, also while Opfield was down, appears asdocker.availability.lease_reacquired. A daemon restart whose copy kept its lease is not a new holding and changes nothing. The best source for the time wins: the new holder’s own report, then the earliest time a voter saw it, and only then the time Opfield noticed. Times a Docker daemon reports are corrected by the measured offset of its clock from Opfield’s when they differ by more than 2 seconds, for example on a virtual machine that resumed with its clock behind; without a measurable clock, the voters’ time or Opfield’s notice decides. When the new holder’s own time is not known yet, for example because it took over while Opfield was down, the entry is written once the time is settled: as soon as the holder reports, otherwise after 60 seconds with the voters’ time, or after 5 minutes when no member reported a time. The details carrytakeoverAt,noticedAt,takeoverSource(holder,voters, ornoticed), andtakeoverId. A waiting entry is stored in the database, so an Opfield restart or crash during the wait loses nothing and never writes it twice. The holder time in the summary (holderSincein the API) is the same takeover time and is corrected as better reports arrive.
Updates and mixed versions
Section titled “Updates and mixed versions”Update Opfield first, then the Node daemons and the Relay Pool; see Relay updates for why the daemons should go first when coming from 2.10. A policy on backend failover stays there while any participant still runs an older release and enters lease mode once, 2 minutes after the last participant was updated. A policy already in lease mode keeps it while only candidate Docker Nodes are outdated; they are excluded and an outdated holder keeps serving until it is updated. An Ingress Node or Relay left on an older release for more than 2 minutes sends the policy to backend failover through Closing, with the serving copies kept running; it returns by itself once they are updated.
Opfield orders daemon updates of lease participants itself. An update of a Docker daemon or Relay that votes in, or can hold, a lease-mode policy waits until the other voters and candidates of that policy are back online, have reported their lease state, and vote again. Meanwhile the Node shows as updating, and the API reports updatePhase: waiting_for_lease_peers with the peers in updateWaitingFor. Updates requested together run standbys first, then holders; Nodes that share no policy update in parallel. A peer that has not settled 3 minutes after its restart stops blocking, and an update that waited 30 minutes fails with an error naming the peers.
The wait matters mostly once. The first restart of a Node updated from 2.10 abstains from voting for about 33 seconds, as does any start after a reboot or with fresh state, so an update may then wait for its lease peers for up to about a minute per peer. Later restarts on the same boot keep voting at once, and a holder whose daemon restarts keeps its slot and keeps serving: its lease is still valid until its watchdog deadline, so its copy takes traffic again as soon as the new daemon process registers, and it keeps renewing until 2 seconds before that deadline. See Docker daemon restarts.
After the update to 2.11, a standby with a managed database binding is prepared once more, because standbys now record the database connection settings they were prepared with.
A Node rolled back to any release since 2.10.0 still starts: daemons keep their state files in a form older releases can read.
Lease details in the API
Section titled “Lease details in the API”get of manage_docker_availability and the Availability API return a lease object:
mode:legacy,bootstrapping,lease, orclosing;reason: why the policy is on backend failover, or why lease mode is impossible right now, withcode,message,nodeIdsorrelayIdsto fix, andsincewhile a lease-mode policy waits out the 2 minutes;excludedNodes:[{ nodeId, reason }];retainedHolderswhileclosing:[{ slot, holderNodeId, confirmed }], the copies that keep running through the close;holders: per slot,holderNodeId,placementId, andholderSince;votersandwitness, withwarningset towitness_near_candidate,no_eligible_witness, orconfigured_witness_unavailable;voterMargin:voters,reachable,required, andmargin. Whenmarginis 0 or less, losing one more voter stops autonomous failover.
Limits
Section titled “Limits”- Data-plane failover protects the workload’s placements, not Opfield, nginx, database engines, registry storage, or shared volumes.
- A copy that cannot renew its lease stops. With Allow available mode off, a network split therefore stops the workload on the side without a majority.
- Planned moves have a short gap for a singleton in Failover mode, as described above.