Relay and Relay Pool
Relay is the authenticated transport that connects Opfield with managed Nodes and carries supported private traffic such as Secure Links. A Relay Pool adds members so transport capacity and availability do not depend on one failure domain. It is not a general VPN or an application load balancer.
Add pool members when the business requires more connection capacity, planned maintenance without losing all transport, or resilience across hosts or sites. The network owner must provide reachable endpoints; the platform owner controls enrollment, assignment, drain, and updates. Success means managed Nodes can establish new sessions through the intended members and private application paths continue through a member outage or drain.
Operator details: topology and endpoint
Section titled “Operator details: topology and endpoint”The local Relay is a required long-lived data-plane service and the sole public owner of 9443/tcp. It authenticates managed-node sessions and supported private tunnel traffic.
A Relay Pool adds supervisor/worker members and fault-domain-aware placement. Operators verify distinct physical domains and advertise reachable worker endpoints; Opfield rebalances assignments automatically, and the Relay Pool update drains, updates, and verifies one member at a time.
Opfield does not create firewall rules, NAT traversal, or a general overlay network. Every assigned managed host must reach its worker data endpoint. During a Relay incident, preserve the identity volume and avoid introducing an alternate unauthenticated path.
Topology
Section titled “Topology”Every installation has a local Relay service associated with Opfield. Additional Relay Nodes run a supervisor that enrolls through the standard Node identity flow and manages a Relay worker. Opfield distributes signed policy and workload assignments; managed nodes establish authenticated outbound connections to their assigned Relay endpoints.
Relay is transport, not workload ownership. Ingress, Docker, database, monitoring, and builder daemons continue to own their host operations. Relay carries authenticated control and supported private streams without becoming a general network tunnel.
Add a Relay Node
Section titled “Add a Relay Node”- Prepare a dedicated host in the intended physical or network failure domain.
- Ensure its advertised address is reachable by participating nodes on TCP
9443. - Open Settings → Relay and choose Add relay node, or create a Relay from Nodes.
- Enter the display name and reachable Relay Address.
- Run the generated one-time installer command on the target host.
- Wait for the enrollment dialog to close after that Node becomes online.
- Confirm the Relay instance reports ready before assigning or rebalancing traffic.
A pending Relay Node is not pool capacity. Do not treat a created database record or installed supervisor as ready until Opfield has verified the worker endpoint and health state.
The installer waits until the supervisor is enrolled and connected to Opfield, and exits with the reason otherwise. One physical host holds one Relay of the pool: enrolling a new Relay on a host that already has one is refused with the name of that Relay, and nothing is enrolled; re-enroll that Relay or remove it first. To change the user of an enrolled Relay, re-run its installer without --token and with GATEWAY_RELAY_RUN_USER: it keeps its identity and takes the Opfield address, certificate pin, advertised address, and port from its configuration. See Run a daemon without root.
Placement and assignment
Section titled “Placement and assignment”Place members in distinct failure domains when resilience matters. Confirm every managed host can reach every Relay that may be assigned to it. Assignment spread controls how many ready Relays receive new workload connections; increasing spread improves redundancy but consumes additional connections and memory.
Nearest Relays first
Section titled “Nearest Relays first”Daemons measure the round trip to every Relay they can reach and report it to Opfield. For each endpoint, Opfield adds the round trip from the endpoint’s Node to the average round trip from the Nodes that connect to it, and makes the nearest Relays the primary ones; the remaining slots of the spread go to standby Relays in other failure domains. Relays within about 20% or 3 ms of the nearest share the primary role, so Relays in one data center share the load. Primaries change only when another Relay is clearly closer, by at least 30% and 5 ms, so round-trip jitter never moves traffic. Until a Node has reported measurements, placement uses a stable hash of the endpoint as before.
Availability members and the Opfield host
Section titled “Availability members and the Opfield host”Two rules apply on top of the spread:
- The Secure Links of Workload Availability placements (member links) go on every ready Relay that supports data-plane failover (
availability_lease_v2), whatever the spread. Standby placements stay registered there too but take no traffic until their Node holds the workload, so after a takeover the new holder is already reachable through every Relay. A new member link starts on those Relays at once. - Every other Secure Link keeps its spread but includes at least one Relay that does not run on the Opfield host, as soon as its Nodes have measured such a Relay as reachable. A link that only the local Relay carries stops when the Opfield host does. When the Nodes of a link can reach no Relay off the Opfield host, the link stays where it works and Settings → Relay shows a warning: “No relay off the Opfield host is reachable from <node>: traffic of this link depends on the Opfield host.” The pool stays healthy; the warning is not an alert.
After an update to a release with these rules, expect one automatic rebalance for each link whose placement changes.
Rebalancing is automatic. When Relay capacity or workload spread changes, Opfield starts a rebalance after the new placement has stayed stable for 30 seconds; existing connections drain without interruption. Before a move, the daemons that use the endpoint probe the new Relay; a probe that the Relay refuses only because the target is an Availability standby counts as verified, and at most 2 probes run per daemon at a time. Automatic rebalancing pauses during a Relay Pool update, but workloads continue to move off drained members.
The pool stays healthy in steady state. Opfield tells three kinds of rebalance errors apart:
- Transient conditions, such as a busy daemon, a Node that is not connected, a restarting local Relay, or a preparation an Opfield restart interrupted, do not count as failures. The move is deferred with the note “Deferred by a transient condition and retried automatically” and tried again after 30 seconds, doubling up to 5 minutes; the pool does not degrade.
- Genuine failures degrade the pool and are retried automatically after 5 minutes; Rebalance retries them immediately.
- A workload removed during a rebalance simply leaves it; the other moves continue.
Relays without Opfield
Section titled “Relays without Opfield”Opfield signs the policy it pushes to each Relay with a lease. A Relay keeps admitting connections on its last signed policy until that lease expires, so Secure Links and database links keep opening new connections through a remote Relay while Opfield is stopped, updated, or unreachable. Set the lease in Settings → Relay under Policy lease: 72 hours by default, from 1 hour to 7 days (168 hours).
Grant lifetime (1–224 hours, 4 by default) sets how long newly issued endpoint and connection grants are valid. Grants of a Relay Pool member always live at least a third longer than the policy lease (96 hours with the defaults), so a Relay never runs out of grants while it still admits on its last policy. A rotation of the grant signing key waits until every Relay has applied the new key or its lease has run out. A Relay from 2.10 keeps a 15-minute lease and caps its grants at 48 hours until it is updated.
What needs Opfield: new and changed Routes and links, revocations, new grants, enrollment, and every management operation. The local Relay stops with the Opfield host, so only Relays on other hosts carry traffic when the whole host is down. When a Relay reconnects after a gap, the audit log records the time it ran on its last policy as relay.instance.policy.stale_period, with from and to.
When Opfield or a Relay comes back, the Relay receives every policy change as soon as it is connected again, and Secure Links that keep serving are not taken apart: if refreshing a link fails because a Node has not reconnected yet, the link keeps its Relay endpoint and Opfield retries.
Registrations and grant changes
Section titled “Registrations and grant changes”- Endpoint generation changes. When an endpoint moves to a new generation, for example after its target’s certificate was rotated or its target moved to another Node, the Relay keeps the previous registration serving, with its tunnels, until the new generation registers. Tunnels are closed only for route or assignment changes and removed endpoints.
- Grants ahead of the policy. A daemon may receive a grant before its Relay has the policy that includes it. The Relay then waits up to 10 seconds for that policy instead of refusing, and a registration refused as “does not match policy” by an older Relay is retried every 0.2–0.4 seconds for the first 10 seconds.
- Unreachable daemons. A Relay drops the connection of a daemon whose host does not acknowledge data within about 2 seconds, usually about 3.5 seconds after the host became unreachable. Its registrations and tunnels end with it, and Availability routes move to other placements. A daemon that is only busy keeps its connection, because its host still acknowledges.
- Restarting daemons. A Docker daemon that restarts gracefully announces it; the Relay keeps its registrations for up to 15 seconds until the next process registers and answers new tunnels with “target endpoint is restarting” meanwhile, which Ingress Nodes hold instead of failing. See Docker daemon restarts.
Route revocation
Section titled “Route revocation”When a Route or link is deleted or its target changes, Opfield pushes a policy that revokes the old route to every Relay. A Relay that has not applied it within 90 seconds is treated as stale for the revoked routes: Opfield drops it from their candidates, and the daemon at the endpoint refuses those routes through it, because every tunnel carries the route it was admitted for. The Relay’s other routes keep working. In Settings → Relay, such a Relay shows “Revocation of N revoked routes pending” and then “Stale for revocation of N revoked routes” until it applies the required policy revision.
Drain, update, and remove
Section titled “Drain, update, and remove”Drain moves new tunnels to other members. Existing streams can continue for up to 10 minutes and are then disconnected; Force disconnect ends the wait immediately. Resume returns the member to service. Drain, resume, and force disconnect require admin:system. A member that is not connected can be drained; the drain applies when it reconnects. Resume and Force disconnect need a connected member and are refused otherwise (409 RELAY_NOT_CONNECTED).
A Relay Pool update performs the rolling procedure itself: it drains each remote member, updates its supervisor and then, through the updated supervisor, its worker, verifies both, and returns the member to service before moving to the next. It waits up to 30 minutes for long-lived streams on a member, such as database connection pools and WebSockets, to end; then it disconnects the remaining ones, records a forced disconnect in the audit log, and continues. Force disconnect ends the wait earlier. A failed run releases the drains it took. A member that is not connected, for example because its host is down, does not hold the update: the run skips it, records Skipped: the relay is not connected; it is updated when it reconnects on its step, updates the other members and the local Relay, and completes. Settings → Relay shows the skipped member, and once it reconnects the Relay Pool update is offered again and updates only the members still behind. The same run updates the local Relay last. When another Relay of the pool is ready and connected and can serve every workload the local Relay carries, and the running local Relay is 2.11.1 or later, the run drains the local Relay the same way first: its workloads move to the other Relays, the drain waits as above, then Opfield recreates the Relay container, verifies that it runs the target version, and returns it to service. The internal registry, which only the local Relay serves, stays on it: a draining Relay keeps admitting image pulls, and a pull cut by the recreate is sent again by the Docker daemon. Otherwise the local Relay is recreated at once, and connections through it drop once for about a second and reconnect: when no other Relay is ready, when the running local Relay is older than 2.11.1 (so the first Relay Pool update to 2.11.1 or later still recreates it at once), when a daemon on a workload’s path has no Relay Pool support, when a workload’s Node reaches no Relay outside the Opfield host, or when the local Relay’s workloads have not moved within 5 minutes of the drain. The local Relay’s step of the run records why. Once the local Relay is being recreated, the run can no longer be abandoned (409 RELAY_UPDATE_COMMITTED) and finishes, or rolls the local Relay back if it fails. In 2.11, Relay releases require Opfield 2.11 or newer, and Opfield does not offer an update while its local Relay is two or more minor versions behind the target.
With a single Relay, recreating it briefly interrupts Secure Link traffic: requests that are open through the old Relay at that moment end, and daemons reconnect within a fraction of a second after the new Relay listens. Ingress Nodes hold new Secure Link connections for up to 3 seconds while the Relay or the target daemon is back, so current daemons see added latency rather than 502 responses. A Relay from 2.10 is stopped after 2 seconds instead of waiting for Docker’s 10-second stop timeout. When updating from 2.10, update the Node daemons before the Relay Pool so the recreated Relay meets daemons that reconnect quickly. Relay updates without any interruption need a second Relay in the pool.
Relays vote in data-plane failover. Before draining a Relay, the update waits until the other voters and candidates of its Availability policies are back and voting, and after the restart it waits until the Relay votes again (at most 3 minutes), so a Relay update never takes a second vote away from a policy.
Remove a remote Relay after it is drained and nothing is assigned to it. An offline member can also be removed, but only after its signed policy has expired and it has not been seen for at least 90 seconds, and only while every affected workload is served by a ready remaining member; Opfield refuses removal that would orphan an endpoint. Until the policy expires, the member could still admit connections on its own. For an offline member, Settings → Relay shows Can be removed after with the date and time in your local time and keeps Remove disabled until then; an earlier removal is refused with 409 RELAY_OFFLINE_REMOVAL_UNSAFE, whose message and details.removableAfter give the same time.
Removing the Opfield record does not replace normal host decommissioning. Uninstall the supervisor and remove host material through the documented operational process.
Relay recovery in 2.11
Section titled “Relay recovery in 2.11”Opfield repairs common Relay trust and certificate problems itself:
- Certificates. Relay server certificates are renewed automatically before they expire; see Certificate renewal. Each member shows whether its certificate is expiring, expired, or failed to renew. Renew certificate appears for an expired certificate or a failed renewal and requires the member to be connected.
- Policy trust, local Relay. If the local Relay rejects Opfield’s signed policy, Opfield re-pins the active signing key automatically, at most once every 10 minutes, and records the recovery in the audit log. A local Relay without this capability shows a message with the manual procedure; update the Relay Pool first.
- Policy trust, remote members. A remote member that no longer trusts any key Opfield can sign with is marked for re-enrollment.
- Re-enroll. For a remote member that needs re-enrollment, has an expired certificate, or is offline or failing, Re-enroll issues a single-use installer command, valid for 7 days, pinned to the Relay version the pool runs. Opfield accepts the token only from the member’s own host. The member keeps its place and assignments; on the host, the supervisor keeps its previous identity aside and restores it if Opfield rejects the enrollment. The installer waits for the enrollment and exits with an error when Opfield refuses the token (already used, expired, issued for another Relay, or run on another host); the Relay then keeps running with its previous identity. To keep a Relay as it is, re-run the installer without
--token. The local Relay cannot be re-enrolled; Opfield recovers it automatically. - Connections after renewal or re-enrollment. Docker and Ingress Nodes follow a member’s renewed or re-enrolled certificate without a daemon restart. Connections that are up stay up through a renewal; a connection that drops is rebuilt with the new certificate and address.
- Lost state. A Relay that restarts without its stored policy, for example after its data file was reset, receives its policy again as soon as it reconnects instead of refusing connections with “policy snapshot is required”.
- Local Relay restarts. With Automatic recovery on, Opfield restarts a local Relay that crashed or hangs. It leaves a restart or stop that someone else started alone, however long it takes, and reconnects to the new Relay within seconds. An operator’s
docker stopof the local Relay is still undone, 2 seconds after the stop finished. A Node whose connection drops during a Relay restart is shown offline only after 5 seconds, and daemons retry registrations with a growing delay of up to 15 seconds, right away once the Relay is reachable again.
Incident checks
Section titled “Incident checks”For a Relay warning, inspect the local Relay or remote worker process, identity volume, PostgreSQL authorization path, advertised endpoint, listener on 9443, policy freshness, and assigned-node reachability. Preserve logs and identity before restart. After recovery, verify node reconnects, Secure Links, database bindings, active streams, assignment state, and the Dashboard warning.
Capacity and resilience validation
Section titled “Capacity and resilience validation”Pool size alone does not prove resilience. Members must occupy independent failure domains and managed nodes must be able to reach the endpoints they may receive. Test connectivity from representative Ingress, Docker, Storage, Monitoring, and Build Worker networks before increasing assignment spread. A second Relay behind the same host, power domain, firewall, or failed route may add capacity without adding availability.
Observe connection count, active streams, memory, reconnect rate, and admission failures during normal load. Establish enough headroom for a remaining member to accept reassigned nodes when one member is drained or unavailable. Rebalancing during an incident can increase reconnect load, so recover the failed dependency first unless assignment concentration is itself the problem.
Recovery and rollback
Section titled “Recovery and rollback”If a member update fails, the Relay Pool update pauses or releases its drains; retry or abandon the run, and restore the signed known-good supervisor and worker release through the supported update path. Do not copy binaries or identity material from another member. Verify its advertised endpoint and readiness before returning it to assignment.
If the pool control state is healthy but one network cannot connect, correct routing, DNS, firewall, or NAT for that advertised endpoint rather than replacing Relay identities. If a remote member’s identity is lost or no longer trusted, use Re-enroll before considering a replacement enrollment, and remove an old record only after its assignments are safe. For a local Relay incident, preserve the configured identity and data volume; recreating the container without them creates a different and unusable transport identity.
After any recovery, test a new managed-node session and a real private path such as a Secure Link or database binding. Existing long-lived streams alone can hide a failure to admit new connections.
For the customer-impact matrix when Opfield, the local Relay, or every Relay is unavailable, continue with Availability, compatibility, and limits.