Skip to content

Hardening checklist

Hardening is the process of reducing avoidable privilege and making recovery predictable before an incident. The security owner approves exceptions, the platform owner implements host and network controls, and service owners prove that tighter controls do not break required workflows. Prioritize controls by customer impact and credential reach rather than by checklist count.

Start with the boundaries that can compromise the entire installation: Opfield host access, administrator identities, encryption and PKI keys, PostgreSQL, Redis, Relay identity, and backup storage. Then harden managed-node roles, workloads, providers, and automation.

  • Keep Opfield and daemon hosts patched and dedicated to their roles where practical.
  • Restrict PostgreSQL and Redis to the private application network.
  • Expose only HTTPS and Relay ports required by the deployment.
  • Require MFA for administrators and minimize administrator membership.
  • Store master keys and backups outside the Opfield host.
  • Use resource-scoped service credentials and short rotation intervals.
  • Prefer immutable image digests and dedicated Build Workers.
  • Deny host bind mounts and privileged/device access unless an explicit supported workflow requires them.
  • Keep managed databases private and use binding identities.
  • Keep TLS certificate verification enabled for external database connections, and enable it for connections that show TLS certificate is not verified after an upgrade to 2.11.
  • Monitor automatic renewal of Opfield, Relay, and daemon certificates, and preserve the system PKI used for internal transport; see Nodes: certificate renewal.
  • Send audit events to an independently controlled SIEM where required.
  • Review source integrations, registry credentials, OAuth clients, tokens, and active sessions regularly.
  • Treat API tokens and OAuth grants that hold administrative scopes, such as admin:system, admin:users, settings:gateway:edit, or nodes:manage, as administrator credentials: since 2.11 programmatic access can carry them. Give each an owner, an expiry, and the narrowest resource scope.
  • Remove temporary maintenance, debug exposure, and manual host changes after incidents.

Do not publish screenshots, support bundles, or logs until secrets, internal addresses, customer names, and repository URLs have been redacted.

  • Run Opfield on a dedicated VM or host and restrict direct shell and Docker access.
  • Keep PostgreSQL, Redis, internal gRPC, and the inference core on private networks.
  • Protect .env, encryption keys, PKI keys, uploaded artifacts, and Relay identity with restrictive ownership and backups.
  • Use HTTPS on the canonical public URL and verify cookie, WebSocket, OAuth, and proxy behavior.
  • Disable unused authentication methods and integrations.
  • Use one enrollment identity per host role and never copy daemon certificates between hosts.
  • Restrict outbound destinations to Opfield/Relay and required providers where practical.
  • Protect systemd units, configuration, update trust anchors, runtime sockets, and local storage.
  • Keep Ingress, Build Worker, Database, and Relay roles isolated according to their risk.
  • Alert on stale last-seen time, capability degradation, disk pressure, and update incompatibility.
  • Prefer digest-pinned images and review source, build logs, vulnerability policy, and provenance boundaries.
  • Run Git builds only on dedicated Build Workers.
  • Use managed volumes instead of host bind mounts.
  • Do not grant privileged mode, arbitrary devices, host networking, or host filesystem access without an explicit supported requirement.
  • Store runtime and Build Secrets through their separate encrypted paths.
  • Keep managed databases unpublished by default.
  • Use one binding identity per application relationship and verify owner credentials never enter workload configuration.
  • Protect Opfield-owned link networks and the secure-link connector from user lifecycle actions.
  • Back up application data through engine-native procedures and test restore independently from Opfield recovery.
  • Keep direct database TLS enabled unless a deliberate private-network exception is documented.

Review administrator membership, group scopes, tokens, OAuth clients, source integrations, DNS connectors, SMTP credentials, SIEM destinations, and active sessions on a fixed schedule. Test denial paths with non-admin identities.

Run incident, restore, update, Relay failover, node restart, workload recreation, and database binding recovery exercises. Remove temporary debug access and document any accepted exception with an owner and expiry.

Some workloads may require a control that the default policy avoids, such as direct database publication, host networking, a device, or broader provider access. Treat that as a reviewed exception, not a hidden toggle. Record the business requirement, affected resources, threat introduced, compensating controls, approving owner, review date, and removal trigger.

Do not weaken a whole Node or installation to accommodate one workload when a dedicated Node, network segment, or external service can isolate the exception. Revalidate exceptions after application, runtime, Node-role, or provider changes.

For each hardening control, prove both the intended operation and the expected denial. Verify administrator recovery after disabling unused authentication, workload recovery after removing host access, build behavior under egress restrictions, certificate renewal with restricted DNS credentials, and database access through binding identities after Node and workload restart.

Keep a recovery method that does not depend on the component being repaired. Backup keys must be accessible when the Opfield host is lost; administrator recovery must work when OIDC or email is unavailable; engine-native backups must be restorable when the control plane is unavailable. Store evidence and review it on a fixed schedule.