Skip to content

Database operations

Database operations combine engine health with the infrastructure that makes the engine usable: Storage Node availability, storage, binding links, workload connectivity, and access identity. Monitor all of those layers before treating a failed application query as an engine incident.

The operating goal is to protect data and restore the customer path with the smallest justified change. The database owner owns backup, integrity, and engine decisions; the platform owner owns Node and storage availability; the application owner verifies real queries and permission boundaries. A successful recovery restores the original client path, explains the failure, preserves evidence, and removes temporary access or exposure.

Agree on recovery point objective (RPO), recovery time objective (RTO), escalation ownership, and restore evidence before the database becomes critical. Monitoring and restart controls reduce diagnosis time, but they do not replace a tested data-recovery plan.

For a running managed database, monitor engine health, Storage Node availability, storage growth, latency, connection pressure, recent operation results, and each binding’s state. The database detail view also exposes the engine-specific metrics, logs, and eligible Explorer or Console functions.

External connections get the same health checks, engine metrics, Explorer, and Console, but no engine logs, lifecycle actions, or bindings, because Opfield does not run the engine. If an external connection shows TLS certificate is not verified, add its CA certificate or use Test and enable verification; see TLS certificate verification. See Managed and external databases at a glance.

A paused instance intentionally disables normal health, metrics, Explorer, and Console behavior until it is unpaused. Do not alert on that expected absence as an engine failure; record the pause window and verify that the service returns to Ready before re-enabling dependent workloads.

Use logs to classify startup, storage, authentication, permission, and recovery errors. Read an operation’s stage before making a corrective change, especially after provisioning, resizing, or deletion. Use Explorer or Console only with the required read, write, or administrative permission, and avoid unbounded production queries through an operator interface.

Explorer browses schemas and rows for PostgreSQL and ClickHouse; schema changes such as adding or altering columns are available for PostgreSQL. Redis uses the command Console. Each Console statement is classified as a read, write, or administrative operation and requires databases:query:read, databases:query:write, or databases:query:admin respectively.

Opfield configuration backups do not replace database data backups. Use an engine-appropriate backup method, retain recovery evidence, and test restores against a separate target. Plan a backup before destructive lifecycle work, substantial resize, or a migration that involves persistent data. The control-plane recovery set is described in Updates, backups, and restore.

On a managed PostgreSQL instance, the database and every application object in it belong to the instance’s application role, and every binding works as that role. The SQL Console runs write and administrative statements as the application role too, so a table created or filled from the Console is usable by every binding at once; read-only statements keep a read-only principal. For a statement that needs Opfield’s own control principal, such as creating a role or an extension that needs a superuser, start the request with RESET ROLE; objects it creates are handed to the application role right afterwards. Opfield also repairs this ownership whenever it creates or reconciles a binding, which fixes tables created from the Console by earlier releases. The repair changes objects only inside the application database and skips system schemas and extension members. Bindings stay ready while the database Node reconnects.

The interactive Console is intentionally bounded by an execution budget shared across statements. It is an operator tool, not a batch migration runner. Prefer versioned migration tooling for schema changes and long-running maintenance, and use a read-only or narrowly privileged identity whenever possible. A browser disconnect must not be treated as proof that a query was cancelled; check the resulting engine state before retrying.

Monitoring collection is background-owned. Opfield starts collection after bootstrap and when a managed database becomes ready; opening the database detail page is not the trigger. On a first visit, a chart may wait for the first real sample or historical query, but the UI must distinguish absent data from a healthy sample.

Check a failed binding in this order:

  1. Confirm that the managed database is Ready and that its Storage Node is online.
  2. Confirm the target Docker node and workload are online and reporting current state.
  3. Review the durable binding desired state and its most recent operation.
  4. Check that the binding’s link on the Node’s shared secure-link connector is reconciled and the workload is attached to the link network.
  5. Verify the binding’s distinct engine principal and intended permissions.
  6. Check the workload’s received configuration and application logs.

Prefer reconciliation or a targeted retry after correcting the root cause. Do not remove Opfield-owned network objects or engine identities manually; doing so can turn a recoverable link issue into a state mismatch or orphaned cleanup problem.

Pause/unpause, restart, resize, credential rotation, certificate rotation, direct publication, binding deletion, and database deletion should each be followed by the relevant verification: engine health, client connection, route or listener state, logs, metrics, and recorded Task result. Direct publication is opt-in; monitor it as an external-client path separately from private bindings.

When deleting a database, remove or migrate bindings first and wait for their link and identity cleanup. A failed delete should remain in Opfield’s operation history for recovery. Do not force-remove storage or users at the daemon or engine layer unless an explicit recovery procedure calls for it.

Define an engine-specific recovery point objective and recovery time objective before production use. On Personal and higher, configure a backup policy with an off-node storage destination and a retention count; otherwise use engine-native tooling. Keep backups off the Storage Node, encrypt them, record the engine version and restore command, and test them on an isolated target. A successful backup without a successful restore test is not recovery evidence.

For a restore exercise:

  1. Create a separate target with compatible engine and storage capacity.
  2. Restore the backup without overwriting the active instance; Opfield restore creates a new managed database by default and requires an empty target.
  3. Run integrity checks and representative application queries.
  4. Recreate access with new bindings or narrowly scoped client identities.
  5. Compare row/key/table counts and application-level invariants.
  6. Record duration, manual steps, and any version constraint discovered.

Opfield does not provide point-in-time recovery merely because it manages the container lifecycle. WAL, Redis persistence, ClickHouse backup strategy, replication, and off-site retention remain database-operations responsibilities.

Classify before changing state: control-plane failures affect Tasks or reconciliation; node failures affect daemon freshness and local runtime; engine failures appear in database logs and health; storage failures affect mount or capacity; binding failures affect listener or principal state; application failures appear after connectivity succeeds. This ordering prevents a healthy engine from being restarted to solve an application permission error.

After recovery, verify the original customer path, not only the administrative screen. Confirm a real query, expected permissions, stable health collection, and that no temporary direct publication, elevated role, debug token, or manual network change remains.