Node roles and installation
These role requirements apply both to nodes enrolled on your own servers and to VMs created through Opfield.
Node roles separate operational authority. The role chosen at enrollment determines which resources the host may control, which software profile is installed, and which failures can affect the business. Choose by responsibility rather than by what packages happen to be present on the machine.
The platform lead should approve the role, host ownership, failure domain, data sensitivity, and recovery owner before enrollment. Success means one durable Node identity is installed on the intended host, reports only the expected capabilities, and passes a low-risk role-specific acceptance check before production resources are assigned.
Operator details: profile prerequisites
Section titled “Operator details: profile prerequisites”Use the generated installer command from the Node creation flow. It contains bounded enrollment material and the pinned Opfield identity.
- nginx: requires nginx, public ingress ports where applicable, configuration validation, and log access.
- Docker: requires a supported Docker Engine and permissions for the configured runtime operations. The installer also installs the lease watchdog,
gateway-lease-watchdog, a separate service that data-plane failover needs on Nodes that hold Availability leases. On a Node installed before the watchdog existed, a Docker daemon that runs as root with systemd or OpenRC installs it itself; for a daemon installed with--user <non-root user>, re-run the Docker node installer withsudoand the same--user, because a non-interactive run otherwise switches the daemon service to root. Each managed disk-image volume holds one loop device, and the daemon installs a boot step,gateway-volume-images, that mounts those volumes before Docker starts. - storage: runs the Docker daemon as root and requires dedicated storage, loop devices, and mount support; the installer formats, mounts, grows, and detaches a test image before enrollment and stops if the host cannot. Each managed database, object storage member, and running backup holds one loop device, so an LXC guest needs its host to pass
/dev/loop-controland a loop-device pool sized for all of them. The Node needs outbound HTTPS toghcr.iofor the backup runner. The earlierdatabasesmode is accepted as an alias. See Storage Node and Docker Node prerequisites. - builder: requires systemd and dedicated BuildKit/containerd capacity, and should not have a Docker Engine socket. On a host without systemd the installer stops before it installs or downloads anything.
- monitoring: provides supported monitoring collection.
- Relay supervisor/worker: extends Relay capacity and resilience with explicit placement.
After installation, validate capability reporting rather than assuming host packages imply support. GPU, Secure Runtime, builder, storage, and migration features depend on reported healthy capabilities.
Choose the role
Section titled “Choose the role”- Ingress: terminates TLS and serves public Domains and Routes. It needs nginx configuration validation, certificate storage, logs, and public ports where applicable.
- Docker: manages ordinary application workloads and Docker resources. It needs a supported Docker Engine and the privileges required for the selected runtime features.
- Build Worker: runs Git builds on isolated BuildKit/containerd services. It must be dedicated to builds and must not expose an ordinary Docker Engine socket to the builder profile.
- Storage: runs managed PostgreSQL, Redis, and ClickHouse, managed SeaweedFS object storage, and database backup jobs with dedicated storage preparation and a restricted Docker daemon profile. Generic application containers and builds are rejected. Nodes enrolled with the earlier Databases role keep their identity, enrollment, configuration, and data; updating their daemon enables the Storage capabilities without re-enrollment or data relocation.
- Monitoring: collects supported host health and metrics without receiving workload lifecycle permissions.
- Relay: installs the Relay supervisor, advertises a reachable address, and joins the Secure Link Relay Pool.
Standard enrollment flow
Section titled “Standard enrollment flow”- Open Nodes and select Add Node.
- Select the intended Node Type and read its description.
- Enter a stable operator-facing name.
- For Relay, enter the address that participating nodes can reach on TCP
9443. - Create the pending Node.
- Copy the generated one-time installation command. It belongs to the Opfield installation that shows it: it downloads the installer from that installation’s release, runs it only after its checksum matches the release, and pins the newest daemon of the same release line.
- Run the command on the intended Linux host.
- Keep the result dialog open while the daemon installs and enrolls; it closes automatically when that specific Node becomes online.
- Open the Node detail and verify role, hostname, version, capabilities, metrics, and service addresses.
The Add relay node action in Relay settings opens the same enrollment flow with Relay preselected and locked. Relay members appear as ready pool capacity only after enrollment and health reporting succeed.
Host preparation
Section titled “Host preparation”Use a dedicated host or VM whenever the role owns a strong trust boundary. Confirm DNS, system time, outbound Relay reachability, package/runtime prerequisites, persistent storage, and systemd or OpenRC before running the installer. Do not copy an installer command between pending Node records.
Storage Nodes require the installer’s storage preflight. Build Workers require dedicated execution and storage capacity. Relay nodes require a reachable advertised address and should be placed in distinct failure domains when the pool is intended to provide resilience.
What the installer checks
Section titled “What the installer checks”- Connection. An installer run, and every re-run on an enrolled Node, succeeds only when the daemon it started runs under systemd or OpenRC and Opfield accepted its session. Otherwise it exits with an error, names the daemon log, and prints its last lines; it waits up to 90 seconds for the session. A token that was already used or has expired is reported as refused by Opfield. Hosts without systemd or OpenRC get a detached launcher that does not survive a reboot.
- Setup command re-run. Re-running a Node’s own setup command, with its already used token, keeps the enrollment and runs as a re-run without a token. A different token on an enrolled host is not used: the Docker and monitoring installers stop with an error and change nothing (on a Node enrolled before 2.11.1 they warn and continue the update), and the nginx installer enrolls with the new token but keeps the previous enrollment when Opfield refuses it. To enroll the host as another Node, stop the daemon, move its
certsdirectory andstate.jsonaside, and run the new command. - Answers. Without
-y, the installer asks on the terminal and stops when it cannot read an answer, for example without a terminal or with its output piped throughtee; a missing answer is never taken as consent. DecliningProceed with installation?or stopping after the summary exits with an error. - Secure Runtime. A fresh Docker Node install sets up Secure Runtime when the host supports it. On a host without support, or when the setup fails, the installer warns with the reason and continues, and the Node reports the same reason in its capabilities. With
--secure-runtimethe install fails instead; the flag also installs Secure Runtime on an existing installation. In an existing/etc/docker/daemon.jsonthe setup adds onlyruntimes.runsc, after saving the original once asdaemon.json.gateway-backup; every other key stays as written. - Daemon version. Without
--version, an installer installs the latest stable daemon release, but a re-run never installs an older daemon than the one installed, for example on a Node running a newer pre-release. It then prints the installed version and the releaselatestresolves to, and keeps the installed version. To install the older release, name it with--version(GATEWAY_NODE_DAEMON_VERSIONfor the daemon installers). This applies to the Docker, Build Worker, Storage, nginx, monitoring, and Relay installers. - Backups. When the monitoring, nginx, and Docker installers replace a file, such as the daemon binary or an nginx configuration, they keep only the newest backup of it (
<file>.backup.<timestamp>).
Run a daemon without root
Section titled “Run a daemon without root”The monitoring, Docker (docker profile), nginx, and Relay daemons can run as an existing non-root user: pass --user <user> to the monitoring, Docker, or nginx installer, or set GATEWAY_RELAY_RUN_USER (and optionally GATEWAY_RELAY_RUN_GROUP) for the Relay installer. The installer itself still runs with sudo and stops before changing the host when the user does not exist. Storage Nodes and Build Workers run only as root.
- Switching user. Re-run the Node’s installer with another
--user, or without--userto return to root. The Node keeps its enrollment and host identity; an enrolled Relay needs no--tokenfor such a re-run. On Docker and nginx Nodes the installer prepares everything while the daemon keeps serving and stops it right before the new process starts; the monitoring daemon and the Relay supervisor are stopped first. The daemon is stopped before any of its files change owner. The switch restarts the daemon once: connections through its database, storage, and container links are cut once for about 1.5 seconds and clients reconnect, as on any daemon restart. Route traffic through Secure Links only waits, because under systemd the daemon of a Docker or nginx Node hands its link sockets to systemd, which keeps them through the switch. On a Docker Node whose storage links still use per-link sidecars from before 2.11.1, the switch also recreates those sidecars. The previous user keeps group memberships the installer gave it, such asdocker; remove them withgpasswd -d <user> docker. - Docker. The installer adds the user to the
dockergroup, which is equivalent to root on the host. Without root, disk-image volumes, moving volume data between Nodes, container log sizes, and installing Secure Runtime from Opfield are unavailable; to install Secure Runtime, re-run the installer with the same--userand--secure-runtime(Node Details shows the command). The lease watchdog still runs as a root service. - nginx. nginx must already run as the same user with
CAP_NET_BIND_SERVICE, and the user must own/etc/nginx,/var/log/nginx, and nginx’s temporary directories; use--nginx-mode integrate. Running nginx as its distribution user (www-dataornginx) and installing the daemon as that user is the simplest setup. Under OpenRC, setcommand_user="<user>:<group>"andcapabilities="^cap_net_bind_service"in/etc/conf.d/nginxandpid /run/nginx/nginx.pid;innginx.conf. When nginx does not run as the user, the installer stops before changing anything and lists what to prepare. To return to root, run nginx as root again first. - Relay. A Relay port below 1024 needs systemd or OpenRC; the installer grants the bind capability.
- Host console.
console.userin the daemon configuration runs console sessions as another user only on a root daemon. A daemon that runs as its own user refuses console sessions whileconsole.usernames another user (409 NODE_CONSOLE_USER_UNAVAILABLE), and the Node page names the fix: removeconsole.useror run the daemon as root.
Alpine, OpenRC, and LXC guests
Section titled “Alpine, OpenRC, and LXC guests”The installers support OpenRC as well as systemd, also for daemons that run as their own user. A Docker Node needs Docker to give its containers the memory, pids, and cpu cgroup controllers, because the secure-link connector and managed workloads run with limits. The installer checks this and refuses to enroll the Node when a controller is missing; it does not change the host’s cgroup setup.
On an Alpine LXC guest with OpenRC, the root cgroup often passes no controllers down, because every process sits in it. Let OpenRC’s cgroups service move the processes out before it enables the controllers: run rc-update add cgroups boot, put the following into /etc/conf.d/cgroups, and reboot the guest.
# Move every process out of the root cgroup so that the cgroups service can pass the controllers down.start_pre() { [ -w /sys/fs/cgroup/cgroup.subtree_control ] || return 0 mkdir -p /sys/fs/cgroup/init for pid in $(cat /sys/fs/cgroup/cgroup.procs); do echo "$pid" > /sys/fs/cgroup/init/cgroup.procs 2>/dev/null done return 0}Then verify with cat /sys/fs/cgroup/cgroup.subtree_control /sys/fs/cgroup/docker/cgroup.controllers: both must list memory, pids, and cpu. The LXC container configuration on the hypervisor needs no change for this.
Verification and cleanup
Section titled “Verification and cleanup”Enrollment is complete only when the daemon reports the expected capabilities and the first inventory/metrics snapshot arrives. If installation fails before enrollment, inspect the local service and installer logs. Delete the pending record only when abandoning that identity; a token from a deleted Node must not be reused.
Installation failure modes
Section titled “Installation failure modes”The nginx node installer can be run again on the same host after a failed run, and a re-run keeps the Nginx mode of the existing installation. If the command cannot download or start the daemon, verify outbound HTTPS, DNS, system time, package-manager state, architecture, and available disk space. If the daemon starts but enrollment does not complete, verify Relay reachability, the one-time token has not already been consumed, and the command was run for the pending Node shown in the dialog. Do not generate several pending records and try their commands interchangeably.
A Node that enrolls without its expected capability is not ready for that role. Check the role service, socket or storage permissions, required kernel features, and the daemon’s bounded startup logs. Correct the host prerequisite and allow capability reporting to refresh; changing the Node Type after installation is not a substitute for installing the correct profile.
If the host was enrolled with the wrong role, decommission that identity deliberately and create a new Node with the correct role. Before deletion, ensure no resources or operations were assigned to the mistaken identity and uninstall its daemon material so it cannot reconnect unexpectedly.
Security after installation
Section titled “Security after installation”Treat the generated installer command as a short-lived credential. Run it only on the intended host, do not paste it into shared chat or persistent shell automation, and remove it from operational notes after enrollment. The installed daemon certificate becomes the durable identity and must not be copied to another machine or restored into two active hosts.
Limit local access to daemon configuration, certificate material, registry credentials, database storage, and role-specific sockets. Enrollment establishes trust between Opfield and the host; it does not harden unrelated services on the operating system. Apply the organization’s host baseline and keep the role’s attack surface no broader than its documented prerequisites.