Skip to content

Architecture

Opfield separates centralized intent and identity from host-local execution.

flowchart TB
    users["Users and automation"] -->|"HTTPS / WebSocket"| gateway["Opfield application"]
    gateway <-->|"Durable state and coordination"| state[("PostgreSQL / Redis")]
    gateway <-->|"Authenticated service channel"| relay["Opfield Relay<br/>TCP 9443"]

    subgraph hosts["Managed hosts"]
        direction LR
        daemons["Role daemons<br/>nginx · Docker · Storage · Build Worker · Monitoring"]
        workloads["Private application and database endpoints"]
        daemons -.->|"Host-local lifecycle"| workloads
    end

    relay <-->|"Outbound control sessions"| daemons
    relay <-->|"Private data-plane streams"| workloads

Private payload moves directly between Relay and host-local application or database endpoints. Opfield authorizes and coordinates those paths, but the Opfield application is not in the payload stream.

The Opfield application is the control plane. It owns users, groups, scopes, resource definitions, desired state, audit records, encrypted integration credentials, long-running operation state, and recovery decisions. PostgreSQL is the durable source of truth for those records. Redis provides required short-lived coordination such as sessions, caches, queues, rate limits, and bounded reservations; it is not a replacement for durable resource state.

Operators and automation submit intent to the control plane through the Operations Console, REST API, OAuth, remote MCP, or AI Workspace. Those entry points share the same backend authorization, validation, plan-entitlement, ownership, and lifecycle checks. Hiding an action in the UI is not the security boundary, and using an API or AI tool does not bypass a product rule.

The control plane does not treat an accepted request as proof that host-local work completed. Mutations that require a daemon are recorded and dispatched as bounded operations. The owning daemon reports progress and observed state, and Opfield exposes durable Tasks, resource status, events, and logs so an operator can verify convergence.

The Relay is a required long-lived data-plane service and the sole public owner of 9443/tcp. Managed daemons connect outbound through authenticated transport; their control interfaces are not normally exposed to the network. The Opfield application keeps its gRPC listener internal and communicates with Relay across an authenticated service boundary.

Relay carries managed-node control traffic and supported private streams. It does not own nginx configuration, containers, databases, builds, or monitoring work on the destination host. It also does not create firewall rules, provide NAT traversal, or act as a general-purpose VPN. Every managed host must be able to reach its assigned Relay endpoint.

If Relay is unavailable, new managed-node operations and new private-link admissions fail closed because identity and authorization cannot be established. An outage of the Opfield application alone is different: a Relay admits connections on a signed policy with a lease, 72 hours by default, so it keeps admitting authorized connections while the Opfield application is stopped or unreachable; see Relays without Opfield. Previously applied workloads may continue locally. Existing data-plane sessions may continue only where the protocol and authorization lifecycle permit it; operators must not infer that a green application process means the control path has recovered.

Each managed daemon has a bounded role and identity. Ingress, Docker, Storage, Build Worker, Monitoring, and Relay profiles expose different capabilities and trust boundaries. A Docker daemon cannot become a database daemon because a client changes a request, and user lifecycle APIs must not mutate Opfield-owned internal containers.

Daemons report capabilities, inventory, health, metrics, service addresses, and operation results. Opfield uses those reports to decide whether a host can safely perform an action. Installed packages alone do not prove capability: the daemon must report the feature as available and healthy. When a node is offline or its capability report is stale, read-only views may remain available from a sanitized snapshot, but mutations that need current host state are rejected.

The daemon owns execution on its host: for example, nginx applies configuration, Docker manages runtime objects, and a Storage Node runs database engines. Opfield owns the durable identity and desired configuration of managed resources. This split lets a workload keep running during a temporary control-plane interruption without allowing the host to invent new authorized state.

Public application traffic is served by Ingress Nodes running nginx. A Domain selects placement, a Route defines traffic behavior, and Opfield distributes certificate material only to Nodes with enabled TLS Routes that require it. The nginx process is therefore part of the data plane, while Domain, Route, certificate, access, and desired-state records remain in the control plane.

Private paths use narrowly defined mechanisms rather than a shared flat network. An nginx-to-workload Secure Link uses an Opfield-owned connector and Relay transport. Managed database bindings, storage links, and container links run through one shared secure-link connector per Docker Node, each link on a private network of its own; a database binding also has a distinct engine identity. These mechanisms have separate ownership and diagnostics even though all provide private connectivity.

Opfield Inference is another separate data plane. A supervised inference core executes provider-facing work and is reached through the Opfield proxy. Published model IDs, provider connections, user limits, accounting, and authorization remain Opfield control-plane concerns.

Use the resource detail page and its durable operation history to determine who owns a failure. A saved desired state with an offline Node points first to host, daemon, time, DNS, certificate identity, or Relay connectivity. A connected Node with an unsuccessful Task points to the role-specific runtime or prerequisite. A healthy runtime with a failing public request points to the Domain, TLS, Route, Secure Link, or application path.

Do not repair an ownership problem by editing implementation children directly. Deployment slots are controlled by their Deployment, Compose-owned containers and networks are controlled by their Compose Project, Route-owned Secure Links follow their Route, and managed database identities follow their binding lifecycle. Direct host changes can create drift that Opfield cannot safely reconcile.

Read views may use sanitized last-known snapshots while a Node is offline. Such data is diagnostic evidence, not current authorization. Mutations remain unavailable whenever Opfield cannot prove ownership, capability, entitlement, or current state.

Desired state and operation records are durable so Opfield can reconcile after application, Relay, daemon, or Node restarts. Recovery should restore the failed owner first, wait for fresh inventory and capability reports, then verify each dependent resource family. Avoid deleting and recreating identities as an initial response: doing so can remove grants, break relationships, or leave the old daemon unable to reconnect safely.

For a path-by-path outage matrix—including the difference between a healthy Relay, a local Relay lost with the Opfield host, and an independent external Relay—see Availability, compatibility, and limits.