Observability overview
Opfield observes control-plane state, node health, workload health, ingress behavior, database health, build activity, inference usage, and security events.
There is no single all-purpose Observability screen. This page describes the operating model across the Dashboard and the detail views for Nodes, Routes, workloads, databases, builds, Tasks, notifications, audit, and status pages. A technical lead should use those signals to answer three questions: are customers affected, which owner must act, and how will recovery be verified?
The desired outcome is not maximum telemetry volume. It is a small set of trustworthy signals with named owners, useful retention, and a tested path from detection to customer-facing recovery. Opfield supplies product and infrastructure signals; service owners still define business-level success, service-level objectives (SLOs), escalation policy, and external dependency monitoring.
Use resource health for immediate state, metrics for trends, logs for diagnosis, durable Tasks for operation progress, notifications for action, status pages for customer communication, audit for attribution, and SIEM for external security analysis.
Design alerts around sustained customer impact and meaningful state transitions. Avoid duplicating every low-level event into a notification channel.
Define ownership and success
Section titled “Define ownership and success”For each important service, name the business owner, technical owner, on-call destination, and customer communication owner. Define what available means from the user’s point of view, not only from an internal process badge. A Route may be healthy while its application or database returns unusable responses; a Node may be online while one customer journey is broken.
Use an SLO—a measurable reliability target over time—to decide which conditions deserve alerts and which belong only in dashboards or reports. Record the expected recovery check in the alert itself. Success is reached when the customer path works again, the signal has recovered, and any temporary mitigation has been removed.
Signal types
Section titled “Signal types”| Signal | Best use | Common mistake |
|---|---|---|
| Health | Current service readiness | Treating one green badge as proof of end-to-end availability |
| Metrics | Capacity and trends | Alerting on every short-lived spike |
| Logs | Detailed diagnosis | Sending secrets or unbounded payloads |
| Tasks | Durable operation progress | Assuming an accepted task already completed |
| Events | Resource state transitions | Using event volume as a health metric |
| Audit | Who changed what | Replacing operational logs with audit records |
| Notifications | Operator action | Forwarding every informational event |
| Status pages | Customer communication | Publishing private topology or raw errors |
Build an operating view
Section titled “Build an operating view”- Start from customer-facing Routes and business services.
- Map each service to its ingress, workload, database, storage, and external dependencies.
- Define the health signal and SLO that represent customer impact.
- Add resource-capacity alerts with sustained windows.
- Route urgent events to an owned notification destination.
- Create a public status component only for information customers should see.
- Test failure and recovery, not only notification delivery.

Diagnosis workflow
Section titled “Diagnosis workflow”Use the Dashboard for triage, then open the affected resource. Compare live health with recent metrics, Tasks, events, and logs. Correlate by stable resource ID, operation ID, request ID, and timestamp. If the displayed state is based on a cached snapshot, restore the owning Node connection before mutating the resource.
For database monitoring, collection starts in the background after Opfield bootstrap and when a managed database becomes ready; opening the database page is not a prerequisite. A first page visit may briefly wait for history or the first real snapshot, but it must not manufacture a healthy state from absent data.
During an incident, begin with the affected customer journey and move inward: public DNS and TLS, Route and Ingress, workload, private dependencies, storage, and external services. Use Tasks to distinguish an operation that was accepted from one that completed. Use audit to understand configuration changes, not as a substitute for runtime logs.
Opfield diagnostics
Section titled “Opfield diagnostics”Opfield reports on its own health as well. The AI assistant and MCP clients read it with manage_gateway_diagnostics:
snapshot: the host’s CPU, memory, load, and data-volume disk; the backend process memory, CPU, and event-loop delay; PostgreSQL and Redis reachability, latency, connection pool, size, and long queries; the stack containers (app,postgres,redis,relay,registry) with state, health, restarts, CPU, and memory; failing background jobs; and API latency and 5xx;history: one-minute samples kept for 48 hours;requestsandjobs: per-route counts, p95, and 5xx, and the last run and errors of every background job;logs: the logs of the app, PostgreSQL, Redis, Relay, and registry containers and of the last Opfield update run.
Reading the state needs diagnostics:view; reading logs needs diagnostics:logs, which includes it. The built-in admin groups hold both; custom groups need them granted. CPU figures cover the CPUs Opfield may actually use, with kernelCpuCount as the kernel’s count. Memory is the container’s cgroup limit when one is set (memoryScope: "cgroup") and otherwise the kernel’s view ("kernel"), as inside an LXC guest, whose own limit is not visible from the app container.
Alert rules in the Opfield category watch the same signals: host CPU, memory, and disk, backend memory, event-loop delay, the API 5xx rate and p95 latency, PostgreSQL latency and pool waits, and Redis latency. They also fire when PostgreSQL or Redis is unavailable, when a stack container is down or unhealthy, and when a background job keeps failing. A PostgreSQL outage alert is sent straight to the webhooks of its rules while PostgreSQL is down and is recorded once it is back. /health checks PostgreSQL over a connection of its own, so a busy connection pool does not read as an outage.
Verify observability itself
Section titled “Verify observability itself”Monitor the monitoring path: node freshness, ClickHouse health, notification queue delivery, status-page publication, SIEM outbox age, and storage retention. An observability system that silently stops collecting must generate a separate platform warning.
Keep telemetry long enough for incident reconstruction, but apply explicit retention and size budgets. Do not use logs as an uncontrolled data archive.
Operator details: stale and conflicting signals
Section titled “Operator details: stale and conflicting signals”Every health decision has a timestamp and an owner. If a Node disconnects, its last snapshot can become stale; restore the owning Node before mutating resources based on old information. If the sidebar warning and a detail page disagree, compare their source, timestamp, and aggregation rule rather than assuming either color is authoritative.
Treat missing data as unknown, not healthy. Verify that collection is progressing, the newest sample advances, ClickHouse is available where structured logs are expected, notification queues drain, and public status updates publish. After repair, run the original end-to-end check and document any blind spot that delayed detection.