Skip to content

Notifications and status pages

Hosting integrations also expose VM power changes, operation failures, synchronization, and supported account-balance thresholds. Billing alerts require billing access and a currency; a running VM is not proof that its Opfield daemon is online.

Create destinations, templates, and alert rules for health, capacity, certificates, deployments, databases, builds, and platform state. Test each destination before relying on it.

Notifications and status pages serve different audiences. Notifications ask an internal owner to act; a status page tells customers what they can expect. The service owner defines impact and priority, the on-call owner receives and resolves alerts, and the communications owner publishes customer-safe updates. Do not send the same raw technical payload to both audiences.

The success criteria are explicit: a meaningful condition opens one actionable alert, recovery closes or resolves it, delivery reaches an owned channel, and a customer-facing incident contains accurate impact without exposing infrastructure details. A configured destination that has never passed an end-to-end test is not production-ready.

Stateful alerts should open when a condition begins and recover when it clears. Configure thresholds and windows to avoid flapping. Include the resource, impact, timestamp, and operator action—not secrets or full daemon errors.

Status pages expose selected service state to customers. Maintenance mode should appear as planned maintenance rather than an unexplained outage. Keep internal topology, private node names, and security-sensitive diagnostics off public pages.

Classify events by customer impact and required action. Capacity trends may start as internal warnings; loss of a public customer journey may require both an urgent alert and a status incident. Avoid creating a public component for every daemon or Node. Components should match products or capabilities customers recognize and should have an owner who can publish updates.

Agree on severity, acknowledgement time, update cadence, and closure criteria before the first incident. Planned maintenance should state the affected capability, expected window, and customer action. Security incidents may require a restricted communication process rather than immediate disclosure of diagnostic detail.

  1. Create the destination and store credentials through the encrypted settings path.
  2. Send a test notification and verify sender identity, signature, rendering, and delivery latency.
  3. Create an alert rule for one meaningful resource condition.
  4. Select a threshold, evaluation window, recovery window, and severity.
  5. Route it to an owned on-call or operational channel.
  6. Trigger the condition in a safe environment and verify both open and recovery messages.

Avoid alerts that operators cannot act on. Every urgent notification should identify the affected resource, customer impact, start time, current state, and the first safe diagnostic action.

Alert rules need notifications:alerts:view to read and notifications:alerts:manage to create, change, or delete them. Webhooks and their delivery history need notifications:webhooks:view, and creating, testing, changing, or deleting webhooks needs notifications:webhooks:manage. notifications:webhooks:manage also reveals webhook URLs, headers, and delivery payloads.

Rules can target Nodes, containers, Git builds, Compose Projects, Routes, Pages, Relays, certificates, databases, and Opfield itself. A container rule on Log Size (MB) finds containers whose logs grow without a limit.

Opfield also raises built-in expiry alerts, naming the owning resource: SSL and Internal PKI certificates at the PKI warning and critical thresholds, CAs well ahead of expiry, Node certificates whose renewal is failing, and Git integration tokens at 30 and 7 days. Each threshold alerts once, and a renewed certificate alerts again in its new lifetime. For managed storage and managed databases, create an alert rule for the Managed Service Certificate Renewal Failed certificate event, which opens while automatic renewal keeps failing and recovers after a successful renewal.

Notification templates use a canonical nested context. Keep templates small and test the exact destination rendering; historical flat variable names are not aliases and can render empty. Treat template changes as operational changes because a broken message can hide the resource, severity, or recovery state even when transport succeeds.

  • Use sustained windows for CPU, memory, disk, latency, and error-rate thresholds.
  • Alert on a state transition rather than repeating the same unchanged state.
  • Separate warning capacity from critical customer impact.
  • Suppress or annotate alerts during explicit maintenance rather than deleting the rule.
  • Review stale, permanently muted, or ownerless rules regularly.

Create components that match customer-visible services, not internal daemon topology. Map incidents to affected components, publish concise updates, and distinguish investigating, identified, monitoring, and resolved states according to the communication process.

Maintenance mode on a managed Route can surface as planned maintenance. Confirm the public status view does not expose node names, private addresses, stack traces, resource IDs, or security details.

Opfield serves the public status page through an Ingress Node, and that Node keeps the last good copy of the page, its assets, and the status data the page polls. While Opfield restarts, is updated, or is unreachable, visitors get that copy, the last one fetched while Opfield answered, so it shows the state from just before the interruption. A page that was never loaded before shows a short “Back in a moment” page that reloads by itself. Browsers never store the page, so live data returns as soon as Opfield answers; while it answers, the data a visitor sees is at most 5 seconds old. By default the status page uses an upstream the Ingress Node can reach. In the status page settings you can also choose an ingress group under Ingress groups (served by every member), so every member serves the page.

Ingress Nodes pick this up when they reconnect to an updated Opfield installation. The first Opfield update to a release with this behavior still interrupts the page, because the cache is empty when the old Opfield version stops; every later update is covered.

Close an incident only after the customer path is verified, not merely when an internal alert turns green. Publish a final summary appropriate to the audience and preserve detailed technical findings in the internal incident record.

Operator details: delivery failure and recovery

Section titled “Operator details: delivery failure and recovery”

If delivery fails, inspect destination health, authentication, TLS validation, response status, retry history, and queue age. Do not rotate a credential without updating every dependent destination. After repair, send a new test and confirm queued delivery behavior before declaring the channel operational.

Maintain a fallback contact path for a failed primary destination. During testing, verify open and recovery messages, deduplication, rendering, links, and timestamps. During a real outage, do not repeatedly recreate destinations or rules: preserve delivery history, repair the narrow failure, and confirm whether queued messages should still be delivered or have become obsolete.