Updates, backups, and restore
Updates and backups protect two different outcomes: updates move the platform forward, while backups make it possible to recover identity, desired state, service credentials, and application data after failure. The platform owner owns the update sequence; the security owner owns recovery keys; database and application owners own engine-native data backups and validation.
Define the recovery point objective (how much data may be lost) and recovery time objective (how long restoration may take) before choosing backup frequency or update windows. Success is a tested restore, not the existence of archive files.
Before an Opfield update, read release notes, verify compatibility floors, back up PostgreSQL and persistent files, preserve encryption/master keys, and record the running image references. Update through the supported installer or generated Compose configuration.
Daemon updates are signed and managed per node. Keep mixed-version behavior within the documented compatibility window and verify reconnect, inventory, bindings, logs, and operations after each role update.
Back up:
- PostgreSQL;
- persistent Opfield artifact/configuration volumes;
- encryption and PKI master keys;
- Relay identity volume;
- engine-native database backups;
- external DNS, identity-provider, email, registry, and SIEM configuration held outside Opfield.
A restore test must prove sign-in, decryption, node reconnection, Relay authorization, Routes, certificates, managed database bindings, and audit continuity. A backup that has never been restored is not a recovery plan.
Database data is outside this recovery set. On Personal and higher, database backups create scheduled native backups of managed and external PostgreSQL, Redis, and ClickHouse databases in a storage connection; on Community, or for engines that need point-in-time recovery, use engine-native tooling. See Managed and external databases at a glance.
Built-in workflows and roadmap
Section titled “Built-in workflows and roadmap”Storage connections and managed object storage and database backup and restore are available in 2.11 on Personal and higher. Opfield configuration export for transfer to another instance is expected in 2.12 on every plan; this is a roadmap estimate, not an available operation today.
Planned configuration export is distinct from existing container archive export. Until it ships, protect the Opfield control plane with the manual recovery set below; database backups protect application data, not Opfield’s own state.
Build the recovery set
Section titled “Build the recovery set”Document the owner, location, retention, encryption method, and restore order for every backup component. Keep master keys and recovery credentials outside the Opfield host and outside the same failure domain as the database backup.
PostgreSQL is the authoritative store for identities, desired state, relationships, operations, and audit metadata. Persistent volumes hold artifacts and service identity material that cannot be reconstructed from PostgreSQL alone. Engine-native managed database backups protect application data; a control-plane backup is not a substitute.
Manual Compose recovery set
Section titled “Manual Compose recovery set”The manual installer uses services named app, relay, registry, postgres, and redis, plus these persistent volumes:
| Volume | Recovery purpose |
|---|---|
postgres_data |
users, resources, desired state, Tasks, and audit metadata |
gateway_data |
uploaded artifacts, generated configuration, TLS, and application state |
gateway_relay_identity |
local Relay identity shared by app and relay |
gateway_relay_state |
local Relay runtime state |
gateway_registry_data |
private registry blobs and manifests |
gateway_registry_auth |
registry token signing material |
redis_data |
persisted Redis state; useful for a coordinated full recovery but not a substitute for PostgreSQL |
Also preserve .env and the exact docker-compose.yml used by the installation; both the installer and the manual installation keep it next to .env, by default in /opt/gateway for the installer. The .env file contains recovery-sensitive secrets; encrypt its backup separately and never attach it to a support ticket.
The following example creates a short offline consistency window. Run it from the directory containing the manual Compose file:
mkdir -p gateway-recovery/{gateway_data,gateway_relay_identity,gateway_relay_state,gateway_registry_data,gateway_registry_auth,redis_data}cp .env docker-compose.yml gateway-recovery/chmod 700 gateway-recoverychmod 600 gateway-recovery/.env
docker compose -f docker-compose.yml stop app relay registrydocker compose -f docker-compose.yml exec -T postgres \ pg_dump -U gateway -d gateway -Fc > gateway-recovery/postgres.dumpdocker compose -f docker-compose.yml stop redisdocker compose -f docker-compose.yml cp --archive app:/var/lib/gateway/. gateway-recovery/gateway_data/docker compose -f docker-compose.yml cp --archive app:/var/lib/gateway-relay/. gateway-recovery/gateway_relay_identity/docker compose -f docker-compose.yml cp --archive relay:/var/lib/gateway-relay/state/. gateway-recovery/gateway_relay_state/docker compose -f docker-compose.yml cp --archive registry:/var/lib/registry/. gateway-recovery/gateway_registry_data/docker compose -f docker-compose.yml cp --archive app:/var/lib/gateway-registry-auth/. gateway-recovery/gateway_registry_auth/docker compose -f docker-compose.yml cp --archive redis:/data/. gateway-recovery/redis_data/docker compose -f docker-compose.yml start redis registry app relayCopy gateway-recovery to encrypted off-host storage and verify its checksum there. This example does not back up databases managed on Storage Nodes; use each engine’s native backup procedure for application data.
To restore into a clean compatible host, first restore .env and the same Compose file, then create stopped containers, copy the persistent data, start PostgreSQL, and import the dump before starting the rest of the control plane:
docker compose -f docker-compose.yml create
docker compose -f docker-compose.yml cp --archive gateway-recovery/gateway_data/. app:/var/lib/gatewaydocker compose -f docker-compose.yml cp --archive gateway-recovery/gateway_relay_identity/. app:/var/lib/gateway-relaydocker compose -f docker-compose.yml cp --archive gateway-recovery/gateway_relay_state/. relay:/var/lib/gateway-relay/statedocker compose -f docker-compose.yml cp --archive gateway-recovery/gateway_registry_data/. registry:/var/lib/registrydocker compose -f docker-compose.yml cp --archive gateway-recovery/gateway_registry_auth/. app:/var/lib/gateway-registry-authdocker compose -f docker-compose.yml cp --archive gateway-recovery/redis_data/. redis:/data
docker compose -f docker-compose.yml up -d postgresdocker compose -f docker-compose.yml cp gateway-recovery/postgres.dump postgres:/tmp/gateway.dumpdocker compose -f docker-compose.yml exec -T postgres \ pg_restore -U gateway -d gateway --clean --if-exists /tmp/gateway.dump
docker compose -f docker-compose.yml up -d redis registry app relayUse the same release first. Verify sign-in, secret decryption, Relay identity, and fresh Node reconnects before attempting an upgrade. File ownership can differ across hosts; if a service reports permission errors, compare ownership inside the original and restored containers rather than applying broad writable permissions.
Opfield update procedure
Section titled “Opfield update procedure”- Read the release notes and required intermediate versions.
- Confirm the active update channel and target version and, for an installation that holds or held a paid plan, that Opfield can reach the license service; see Update prerequisites in 2.11. When updating from 2.10, follow Updating from 2.10 first.
- Verify recent PostgreSQL and persistent-volume backups.
- Record current Opfield, Relay, and daemon versions and image digests.
- Confirm PostgreSQL, Redis, Relay, and disk health.
- Place customer-facing services into maintenance only when the update requires it.
- Apply the supported update without replacing persistent volumes or master keys.
- Wait for Opfield readiness and schema migration completion.
- Verify Relay, sign-in, Dashboard bootstrap, node reconnects, Routes, and database bindings.
- Remove maintenance only after external verification.
Do not treat a running container as a successful update. The application must decrypt existing secrets, read prior state, reconcile background services, and expose healthy customer paths.
Update prerequisites in 2.11
Section titled “Update prerequisites in 2.11”Every Opfield update first authorizes the target release with the license service at license.thesqlabs.com. This applies to updates started from the Console and to re-running the installer. Opfield verifies the signed release before it changes the host:
-
Paid installations: an installation that holds or held a paid activation—including an expired, revoked, transferred, or deactivated one—receives the commercial core for the target version, so its existing paid resources keep running after the update. Opfield verifies the core’s signature, version, sizes, and hashes, and accepts the authorization only as a signed license state for that request. Such an installation needs the license service during the update: if the service cannot be reached, refuses the release, or returns a state that fails verification, the update stops before any change and the current Opfield keeps running. Opfield never silently replaces a paid installation with Community, so allow outbound access to the license service during update windows.
-
Community installations: an installation that never held a paid plan updates as Community. When it has no license key, no paid plan in its license state, and no commercial core on the host, it also updates while the license service cannot be reached, because it downloads nothing private. A refusal or a license state that fails verification still stops the update.
-
Running operations: after pulling images and before changing the host, Opfield waits for running and queued orchestration work—blue/green deploys and the drains of their previous slots, Compose operations, Workload Availability actions, migrations, and build rollouts—to finish. The wait lasts up to 15 minutes by default, or up to the longest deadline announced by running work, and at most 60 minutes. The update screen lists the operations it waits for. New orchestration work started during the wait is rejected with
GATEWAY_UPDATING. An administrator can end the wait early with Update now; Opfield then resumes or reconciles the interrupted operations after the restart. -
Free disk space: before it replaces the app, the update takes a snapshot of the Opfield database and needs free space for it on the filesystem that holds the Opfield directory: the database size plus 10% plus 256 MiB. With less, the update stops before any change, and the update container logs say how much space is needed.
The previous image and commercial core are kept so that a failed update can roll back. See Commercial core for how paid features are delivered.
The wait for running operations, the disk-space check, and the database snapshot below belong to the updater of the version that is running. They apply to updates started from 2.11 on, not to the update from 2.10; see Updating from 2.10.
Automatic rollback and the database snapshot
Section titled “Automatic rollback and the database snapshot”Just before the new version starts, the update writes a pg_dump custom-format snapshot of the Opfield database to .gateway-foundation-backups/pre-update-<time>/gateway-db.dump in the Opfield directory, next to the saved .env and docker-compose.yml. If the new version does not become healthy within about five minutes, the update container rolls back: it restores .env and docker-compose.yml, stops every Compose service except postgres, drops the gateway database, and recreates it from the snapshot. The previous version therefore never runs on a schema changed by the new version’s migrations. Data written after the snapshot is lost.
The snapshot is deleted once the update or the restore succeeds, and a later update deletes any leftover snapshot older than 7 days. If the migrated database cannot be dropped, the previous version starts on it as before and the snapshot is kept. If the restore fails after the drop, Opfield is left stopped and the snapshot is kept: fix the cause shown in the update container logs, then restore from the Opfield directory:
docker compose exec -T postgres psql -U gateway -d postgres -c 'DROP DATABASE IF EXISTS gateway WITH (FORCE)'docker compose exec -T postgres pg_restore --create --exit-on-error -U gateway -d postgres < .gateway-foundation-backups/pre-update-<time>/gateway-db.dumpdocker compose up -dThe snapshot protects only the update itself. Keep the PostgreSQL backups of the recovery set as well.
Updating from 2.10
Section titled “Updating from 2.10”The update from 2.10.x is run by the 2.10.x updater. It neither waits for running operations nor snapshots the Opfield database, so prepare it yourself.
Before you update
Section titled “Before you update”-
In the Opfield directory (
/opt/gatewayby default), back up the database and keep a copy of.env, which holdsPKI_MASTER_KEY:Terminal window docker compose exec -T postgres pg_dump -Fc -U gateway gateway > gateway-2.10.dumpcp .env env-2.10.backup -
Avoid deploys, backups, and Docker migrations while the update runs.
-
On a paid installation, make sure the Opfield host can reach the license service: the update downloads the signed commercial core of 2.11 and stops before the app is replaced if it cannot. Community installations that never held a paid plan update without it.
-
On Community, read Upgrading a Community installation from 2.10: from 2.11, databases (including external connections saved in 2.10), storage, GitLab integration, and AI Workspace Plan Mode, Scenarios, and Sandboxes need Personal or higher. Their records stay in the database and become usable again with a Personal or higher key.
The update applies the 2.11 database migrations. Existing data that 2.11 no longer allows, such as two workloads on one host port of a Node, two enabled Routes with the same name on a Node, or duplicate active backup runs, is flagged in the audit log instead of failing the update; see Conflicts resolved by the 2.11 update.
If the update rolls back, 2.10.x starts again on the already migrated database and keeps its permissions and a valid license; updating again restores everything 2.11 had. Fix the cause shown in the update container logs and update again. To return to the state before the update instead, put the saved .env back, stop every service except postgres (docker compose stop app relay registry redis), restore gateway-2.10.dump with the DROP DATABASE and pg_restore commands from Automatic rollback and the database snapshot, and run docker compose up -d.
Update order
Section titled “Update order”- Update Opfield.
- Update the Node daemons with Nodes > Update Nodes, which lists every Docker, Nginx, Storage, and Monitoring Node that has a newer daemon and updates the selected ones together.
- Update the Relay Pool with Settings > General > Update Relay Pool.
Do not update Node daemons or Relays before Opfield: the 2.11 Relay requires Opfield 2.11, and many 2.11 features and fixes need the 2.11 Docker, Nginx, and Relay releases. Relays from 2.10 keep the 15-minute relay policy lease until they are updated.
After the daemon and Relay updates:
- The first apply of a Compose project, even of an unchanged revision, recreates its services once, because they receive the new log settings and a label with their own configuration digest; later revisions recreate only the services they change; see Supported Compose configuration. Deployment routers stop capping request bodies at their next deploy, slot switch, or restart.
- Docker Availability policies switch to lease mode by themselves once every Docker Node, Ingress Node, and Relay of the workload has run 2.11 for 2 minutes. Each Docker Node needs the lease watchdog: the Docker node installer installs it, and a 2.11 Docker daemon that runs as root with systemd or OpenRC installs it itself. Otherwise the policy lists the Node under Excluded nodes; re-run the Docker node installer there with
sudoand the same--user. In lease mode every Node of the policy must reach every Relay. See Data-plane failover. - Storage Nodes need outbound HTTPS to
ghcr.ioto pull the backup runner. Other third-party runtime images come from the Opfield mirror onghcr.iofirst and fall back to Docker Hub. - External PostgreSQL and Redis connections with TLS created before 2.11 stay unverified and show a warning. Run the one-click check or add the private CA; see TLS certificate verification.
- Claude Fable 5.1 and Opus 5.5 models published in Opfield Inference before 2.11 have tools disabled; open and save them again.
Permissions after the update
Section titled “Permissions after the update”- Groups, users, API tokens, and OAuth grants that held a retired scope name received its replacement, and retired names are still accepted on input for two releases. Scripts that read scope names back should expect the new names. See Retired scope names.
- Some replacements are broader than the old scope. Review groups, user permissions, API tokens, and OAuth grants that held any of these:
integrations:gitlab:ci:edit,:variables:edit,:variables:delete,:webhooks:manage, and:registry:managebecameintegrations:gitlab:repo:write, which also commits repository files and changes CI configuration, CI/CD variables, webhooks, and registry settings. A CI token that only cleaned up the registry can now push commits.integrations:gitlab:ci:viewbecameintegrations:gitlab:repo:read, which also reads repository files.integrations:gitlab:sync,:github:sync, and:git:syncbecame the connector’smanagepermission: settings, tokens, allowlist, test, and delete.nodes:config:editbecamenodes:manage, which also installs the Secure Runtime, changes service addresses, and powers, resizes, and restores snapshots of hosted VMs.proxy:advanced:bypassandproxy:raw:bypassboth becameproxy:unrestricted, so either one now lifts both restrictions.proxy:templates:create,:edit, and:deletebecameproxy:templates:manage, every nginx template operation.notifications:alerts:create,:edit, and:deletebecamenotifications:alerts:manage, andnotifications:webhooks:create,:edit, and:deletebecamenotifications:webhooks:manage, which also reveals webhook URLs, headers, and delivery payloads.docker:containers:folders:managebecamedocker:folders:manage, which covers the folders of every Docker resource type.
integrations:gitlab:variables:viewbecameintegrations:gitlab:repo:readand no longer reads CI/CD variable values; reading values needsintegrations:gitlab:repo:write.- Every action permission except create now implies the view permission of its family with the same qualifier: a group or token that held only a delete, edit, manage, or run permission can now also list and view those resources. Review delete-only grants.
- Built-in groups receive the new permissions automatically. Custom groups do not: grant
diagnostics:viewanddiagnostics:logsfor Opfield diagnostics,databases:backups:view,:manage,:run, and:restore,nodes:backups:execute, andstorage:view,:create,:edit,:delete,:iam,storage:objects:*, andstorage:credentials:usewhere needed.storage:credentials:revealimpliesstorage:credentials:use. - Tokens never see more than the account that minted them, and a create permission no longer counts as view when access is delegated: tokens of users whose access shrank lose that access.
- Git sources saved after the update build automatically only while the account that saved them keeps
useon the repository; webhooks and automatic deploys check the account that configured them. - Deploying or updating a Deployment with a new image, command, or environment, and updating a container’s image, need the environment and secrets permissions; attaching a container to networks needs network edit access. Write-scoped SQL on an external PostgreSQL runs one statement at a time, and schema or role changes need the admin permission.
- Community plan limits are 25 managed Nodes, 3 users, and 1 custom permission group (100, 10, and 5 in 2.10). Nothing existing is removed; creating more is refused.
Changed behaviour
Section titled “Changed behaviour”- TLS for unknown names. Ingress Nodes reject TLS handshakes for hostnames that no Route serves. Clients or load balancers that connect by IP over HTTPS without SNI, or rely on a default site, are refused. Nodes whose nginx already has its own default server on port 443 skip this with a warning. See Routes.
- Compose restrictions. Compose projects cannot use the host network or Opfield’s internal networks (
gateway-secure-links,gateway-db-*,gateway-storage-*), set Opfield’s reserved labels, or set Docker client variables such asDOCKER_HOST,DOCKER_CONFIG,BUILDKIT_*, andCOMPOSE_*. Existing projects that break these rules keep running and can be stopped and removed, but a new revision, apply, start, or restart is refused until they are edited, including the first apply after the Docker daemon update. See Supported Compose configuration. - nginx directive checks. nginx template content and Additional Route advanced configuration are checked like Route configuration: without
proxy:unrestricted, denied directives are refused and nginx logs can be written only under/var/log/nginx. Templates saved before the update keep rendering as saved. - Relay policy lease. Relays keep working without Opfield for 72 hours by default (Settings > Relay, 1 hour to 7 days); see Relays without Opfield.
- Housekeeping. Operation history—finished builds with their logs, Compose, Availability, and hosting operations, and webhook deliveries—is kept for 90 days, and expired OAuth grants and OAuth clients that never completed authorization are purged. Change Settings > Features > Housekeeping before the first run if you need a longer history; see Housekeeping and retention.
- Internal PKI renewal. Internal PKI certificates linked to Routes are reissued automatically before expiry when Opfield holds the private key; auto-renew is turned on for existing links.
- GitLab token rotation. GitLab tokens with the
apiorself_rotatescope that expire within 14 days are rotated by Opfield, which revokes the old token. Give Opfield a token of its own if the same token is used elsewhere; see Token expiry and rotation. - Settings. Saving settings restarts Opfield only when internal HTTPS is switched, or when the Public URL changes while internal HTTPS is on. Sign-in methods, MFA, the OIDC provider, and identity provisioning moved to Settings > Authentication, structured logging is on Settings > Features, and Authentication email (SMTP) is now SMTP configuration.
- API. Node reads no longer include enrollment token material, an empty Node update returns
400, and a malformed ID returns404. Environment, recreation, and Compose revision changes return409while a build rollout owns the target. An unknown/apipath returns a JSON404. - Docker Availability is no longer labelled Tech Preview, and builds can no longer be pinned to the Dashboard or sidebar.
Conflicts resolved by the 2.11 update
Section titled “Conflicts resolved by the 2.11 update”Opfield 2.11 enforces several uniqueness rules in its database rather than only in application code, so they also hold when several Opfield processes handle requests at once. The update never fails on existing data that breaks one of them: migration 0207 keeps the earliest record, flags or renames the others, and writes every case to the audit log. After the update, search the audit log for these actions and resolve what they report.
| Rule after the update | What the update did with existing conflicts | Audit action |
|---|---|---|
Two enabled Routes on one Ingress Node cannot serve the same name (409 PROXY_HOST_DOMAIN_CONFLICT) |
The earliest Route keeps the name. Later ones keep their configuration and stay flagged until their names or node change. Once disabled, such a Route cannot be enabled again while another enabled Route serves the name. nginx serves only one of them, so remove the name from one Route or disable it | proxy_host.domain_conflict |
A backup policy has one queued or running backup (409 BACKUP_ALREADY_RUNNING), and one restore into a given new database name runs at a time (409 BACKUP_RESTORE_ALREADY_RUNNING) |
One active run continues: a running one over a queued one, otherwise the newest. The others are marked failed with the phase superseded, release their executor, and have their runners cancelled |
database.backup.superseded |
Managed storage names are unique per Storage Node (409 MANAGED_STORAGE_NAME_IN_USE) |
The earliest cluster keeps the name. Later ones, and their storage connections, are renamed with the first free -2, -3, … suffix |
storage.managed.renamed_duplicate |
A host port on a Node belongs to one Deployment, managed storage, or managed database (409 DEPLOYMENT_HOST_PORT_IN_USE, MANAGED_STORAGE_PORT_CONFLICT, or MANAGED_DATABASE_PORT_CONFLICT) |
The earliest workload keeps the port. Later ones keep running and are flagged; move one of them to another port | node.host_port_conflict |
A certificate that a Route uses cannot be deleted (409 CERT_IN_USE) |
Nothing to resolve: the database now enforces the existing check | — |
Host ports that the Docker daemon picks for a managed database avoid the ports other workloads reserve on that Node. A managed database whose published port still conflicts with another workload’s reservation lists it in managed.hostPortConflicts of its database connection, with the port and the workload that holds it.
The update does not delete Routes, workloads, storage, or backups, and does not stop running workloads. Clients that rename managed storage or look it up by name should be checked for the renamed clusters.
Updating to 2.11.1
Section titled “Updating to 2.11.1”2.11.1 moves database bindings and storage links to one shared secure-link connector per Docker Node and adds container links. Update Opfield first, then the Docker daemons. What happens on each Node after its Docker daemon update:
- Each workload with a database binding is recreated once, one at a time on each Node, the next once the previous runs healthy. Deployments roll out blue/green without downtime; standalone Containers and Compose services restart once, so schedule the update of Nodes that run such workloads accordingly.
- Each workload with a storage link created before 2.11.1 is recreated once right after its link switches. Links created on 2.11.1 need no recreate.
- Variable names and credentials do not change, and nothing needs to be done by hand. Do not delete old link networks or containers yourself.
- Rolling a Docker daemon back to 2.11.0 switches its links back automatically, with one recreate per workload. At most four workloads per Node are recreated at once, the daemon’s limit on concurrent commands, so on a Node with many linked Deployments the later ones reconnect a little later.
- Later connector updates keep open link connections for up to 30 minutes, then close them once.
New link networks use their own address range, 10.213.x.x (a /16) by default. If it overlaps a network your Nodes reach through their default route, set docker.secure_links.subnet_pool before updating the Docker daemon, so the moved links already use your range (earlier daemons ignore the key); see Shared connector and link networks.
Daemon and Relay updates
Section titled “Daemon and Relay updates”Update one failure domain at a time. For Relay Pool members, drain, update, verify, and return each member before moving to the next. For ordinary Nodes, verify role-specific resources after reconnect. Do not update every Ingress, Storage, or Relay member simultaneously unless the environment has an independently tested recovery path.
Nodes > Update Nodes updates the daemons of several Nodes together, and Opfield refuses daemon updates for disconnected Nodes, Relay Nodes, and Nodes that already run the release; see Daemon updates. After an Opfield update, update the Node daemons before the Relay Pool, especially when coming from 2.10: the Relay Pool update also updates the local Relay, recreating it at once while it runs a release older than 2.11.1, and current daemons reconnect to it within a fraction of a second; see Relay and Relay Pool. Daemon and Relay updates of Nodes that take part in data-plane failover are sequenced by Opfield, one member of each Availability policy at a time, so an update can wait for its lease peers.
Restore procedure
Section titled “Restore procedure”- Provision a clean compatible Opfield environment.
- Restore PostgreSQL and persistent volumes from the same recovery point.
- Restore encryption, PKI, and Relay identity material with original permissions.
- Start internal dependencies, Opfield, and Relay in the documented order.
- Verify sign-in and decryption before permitting mutations.
- Allow managed nodes to reconnect with their existing identities.
- Reconcile Routes, certificates, workloads, database bindings, notifications, and integrations.
- Restore application databases through engine-native procedures where required.
- Confirm audit continuity and record the recovery event.
If only part of the recovery set is available, stop and assess the consequence. Generating replacement master keys or identities can permanently orphan encrypted state or managed nodes.
Rollback decision
Section titled “Rollback decision”Decide the rollback trigger before starting an update: failed migration, inability to decrypt existing secrets, Relay authorization failure, incompatible daemon state, broken customer Routes, or another measurable condition. Keep the prior approved image references and the matching backup until the new release completes its observation window.
Application rollback and data rollback are separate. Apart from the automatic rollback of an update that never became healthy, reverting an Opfield image does not reverse a completed PostgreSQL schema migration, and moving a workload back to an older artifact does not revert mutable volume or database data. Follow the release-specific compatibility guidance and restore data only from a coordinated recovery point when integrity requires it.
Operator validation after update
Section titled “Operator validation after update”| Layer | Required proof |
|---|---|
| Identity | Existing users sign in; MFA, OAuth, and decryption work |
| Control plane | Migrations complete; background schedulers and audit continue |
| Relay and Nodes | Versions are compatible; Nodes reconnect with fresh inventory |
| Customer traffic | External DNS, TLS, Route health, and application response succeed |
| Data | Managed databases are healthy and application bindings can query |
| Automation | Builds, webhooks, notifications, and source connectors used by the installation work |
Keep the update record with start and end time, versions, image digests, backup identifiers, operator, exceptions, and final verification. This makes the next update a controlled procedure instead of a reconstruction from shell history.