Availability, compatibility, and limits
Opfield separates management availability from service availability. A control-plane outage removes the ability to make and authorize new changes, but it should not stop infrastructure already applied to managed hosts.
This distinction is central to production planning: the customer-facing application path, the Opfield application, and Relay may have different failure domains.
What continues during an outage
Section titled “What continues during an outage”| Failure | What normally continues | What becomes unavailable or uncertain |
|---|---|---|
| Opfield application unavailable; Relay healthy | nginx keeps serving applied configuration; containers and databases keep running; Relays keep admitting new Secure Link and database-link connections on their last signed policy for the policy lease (72 hours by default); Availability policies in lease mode keep failing over; the public status page is served from the Ingress Node’s cache | UI, API, desired-state changes, new or changed Routes and links, revocations, new managed operations, and backend failover of Availability |
| Opfield host unavailable; local Relay runs on the same host | public routes and workloads that do not depend on Relay continue on their own hosts; links that also use a Relay on another host keep working through it, and lease-mode Availability fails over through such Relays | local Relay also fails; managed-node control is unavailable, and links carried only by the local Relay break |
| One external Relay member unavailable | workloads continue; assigned traffic can use other ready members only where placement and active assignments provide that path | sessions assigned only to the failed member may reconnect or fail; new capacity is reduced until recovery and rebalance |
| No reachable Relay | ordinary host-local workloads and directly served public traffic may continue | managed-node control and new private links fail; existing Relay-dependent streams are not guaranteed |
| One managed Node unavailable | resources on other Nodes continue; an Availability policy in lease mode moves the Node’s slots to other candidates within about 45 seconds, with or without Opfield | resources owned by that Node are unavailable or stale; Opfield must not mutate from stale inventory; on backend failover, replacements need a running Opfield |
| License key ordinarily expires, or a paid plan is downgraded, revoked, or deactivated | everything during the grace period (none for revocation, transfer, or deactivation); after it, existing workloads, the data plane, data, scheduled backups, and viewing, monitoring, and deletion of existing paid resources | creating paid resources and changing their configuration; SIEM forwarding and external registry tokens pause until renewal; see Downgrade and expiry |
| License service unreachable | Opfield, workloads, and paid features continue; a previously valid paid key stays valid for 100 days after its last signed license state | activating a key and updating an installation that holds or held a paid plan (a Community installation without a key still updates); after 100 days, the after-grace rules of an ordinary expiry |
| Valid paid key, commercial core missing or rejected | existing infrastructure and Community features continue | paid features are reported as not ready until Enable paid features succeeds |
“Continues” means the last successfully applied local state remains active. It does not guarantee that an application has enough replicas, that its database is healthy, or that an existing private stream survives every network failure.
Workload Availability
Section titled “Workload Availability”Workload Availability (HA) is available on Business and Enterprise. It supports 2–32 replicas or one serving placement with replacement for mount-free Containers, Deployments, and whole Compose Projects. When every Docker Node, Ingress Node, and Relay of a workload runs a release with data-plane failover, the Nodes replace a lost holder themselves, even while Opfield is down, and a cut-off Node stops its own copy. Otherwise Opfield restores the requested placements on available eligible nodes, which requires a healthy control plane. Both need capacity, artifacts, and dependencies; neither provides HA for Opfield itself, nginx, databases, or storage.
The local Relay concern
Section titled “The local Relay concern”Every installation includes a local Relay. If Opfield and that Relay share one host, losing the host removes both management and the transport used by Secure Links and managed-node sessions.
For environments that require private connections to survive loss of the Opfield host, add external Relay members in independent failure domains and verify that every relevant Node can reach them. Opfield then places every Secure Link on at least one Relay off the Opfield host that its Nodes can reach, and Settings → Relay warns about links that depend on the Opfield host because their Nodes can reach no other Relay. Merely running a second Relay container on the same machine, firewall, power source, or network path does not improve fault tolerance.
Test the exact application path. A public Route to a local upstream may continue without Relay, while a Route or application database binding that uses a Secure Link depends on reachable Relay transport.
Version compatibility
Section titled “Version compatibility”Treat Opfield, Relay, managed daemons, and their protocol versions as one tested release set.
- Use the versions published together by the release metadata and supported installer.
- Update Relay Pool members one failure domain at a time.
- Update representative Nodes first and verify fresh capabilities, inventory, and a real operation.
- Do not assume arbitrary older or newer daemons are compatible because they can establish a TCP connection.
- Keep the previous approved image references and matching backup until the observation window ends.
- Do not downgrade across a database migration unless the release notes explicitly support it and a coordinated restore point exists.
Data-plane failover needs the same release on every participant of a workload: a policy enters it only after all of them have run it for 2 minutes and leaves it when an Ingress Node or Relay stays outdated for 2 minutes. See Updates and mixed versions.
Opfield does not currently publish a universal long-term compatibility matrix for every historical component combination. If your policy requires an extended mixed-version window, validate that exact combination during the pilot and record it as an installation-specific constraint.
Capacity and tested limits
Section titled “Capacity and tested limits”There is no single meaningful maximum number of Nodes, Routes, workloads, builds, databases, or concurrent private streams. Capacity depends on host resources, operation rate, retention, payload shape, Relay placement, build workload, and external providers.
Do not convert plan quotas into a performance guarantee. Before production, test a representative workload at expected peak plus agreed headroom and record:
- Opfield API latency and background queue delay;
- PostgreSQL connections, latency, storage growth, and backup time;
- Redis memory and queue health;
- Relay connections, memory, throughput, reconnect behavior, and assignment spread;
- Node report volume and freshness;
- build concurrency, artifact size, and disk pressure;
- log, audit, metrics, and Pages retention growth;
- recovery time after restarting each shared component.
If a required limit or SLA is not published for your release, treat it as not guaranteed. Establish it through a pilot, keep the evidence with the named release and topology, and retest after material upgrades.
Production decision
Section titled “Production decision”Before accepting a workload, document which outage paths it tolerates, which paths require external Relay, which data services have independent high availability, and who restores the control plane. Then run the production checklist and incident runbook against that topology.