Skip to content

Ingress troubleshooting

The incident objective is to restore the public service without destroying the state that explains the failure. Ingress crosses several ownership boundaries—DNS, network, certificate, nginx configuration, Route policy, private transport, and application runtime—so changing several layers at once usually increases outage time.

Assign one incident owner, record the last known-good change, and verify from the outside in. Success means external behavior is restored, Opfield desired and reported state agree, and the failed layer and recovery action are recorded well enough to prevent recurrence.

Use this order to avoid changing the wrong layer:

  1. DNS: resolve the hostname from an external client and confirm it points to the assigned ingress node.
  2. Network: confirm public 80/443 reachability and host firewall rules.
  3. TLS: inspect certificate hostname, chain, expiry, and distribution status.
  4. nginx config: verify the latest revision applied and configuration validation succeeded.
  5. Route health: inspect expected status/body and maintenance state.
  6. Upstream: test the application from the nginx node or inspect Secure Link health.
  7. Logs: correlate nginx access/error logs with workload logs and request IDs.

Common mistakes include cross-node Domain/Route placement, stale external DNS after migration, HTTP-01 on a node without public port 80, a WebSocket application without WebSocket forwarding, and selecting a workload port that is not actually published or linked.

Symptom Most likely layer First evidence
Hostname does not resolve DNS External A/AAAA lookup and Domain placement
Connection timeout network Public address, firewall, ports 80/443, node availability
Certificate warning TLS Certificate hostname, chain, expiry, and assigned Route
TLS handshake fails at once Route matching No enabled Route on that node serves the hostname; the node rejects such handshakes instead of answering with another Route’s site
Immediate 404 Route matching Hostname, path prefix, enabled state, raw config
Managed 503 maintenance or unhealthy upstream Maintenance state and health history
502/504 upstream transport Secure Link, target port, workload health, timeout settings
WebSocket disconnects protocol forwarding WebSocket option, application path, proxy/read timeouts
Changes never appear apply/reconciliation task state, nginx revision, node connection, validation error; a rejected change fails with 422 NGINX_CONFIG_FAILED and reports the nginx -t output
502 from a Docker Route in Raw Config Mode raw configuration The raw configuration must keep proxying to the Route’s Secure Link upstream; see Routes
502 from a Deployment Deployment router Header X-Gateway-Deployment-Router: upstream-unavailable means the router could not reach the active slot

Record the failing hostname, path, time, client IP class, expected response, actual response, assigned node, Route ID, latest task ID, and relevant request ID. Preserve the generated-config validation error and bounded logs before retrying; repeated edits can erase the most useful evidence.

Compare three views:

  1. Opfield desired state in the Route and Domain details;
  2. latest acknowledged state from the nginx node;
  3. externally observed DNS, TLS, and HTTP behavior.

A mismatch between these views identifies whether the failure is control-plane state, node apply, or external infrastructure.

Secure Link failures are logged once per link, not once per request. The nginx daemon on the Ingress Node logs a warning proxy secure-link connections failing with the link_id, the last Relay, stage, and error when a link starts failing, still failing with a count every 5 minutes, an info line when requests succeed only after retries, and recovered with the numbers of failed and retried requests at the first clean connection. The Docker daemon of the target logs failing incoming tunnels the same way per endpoint owner (relay endpoint tunnels), and refusals on revoked routes separately (relay endpoint tunnels on revoked routes). Per-attempt details are at debug level.

  • Correct DNS only after the intended node placement is confirmed.
  • Correct a certificate assignment rather than disabling TLS globally.
  • Fix config validation before forcing another apply.
  • Use maintenance mode when the application must remain unavailable during repair.
  • Retry reconciliation only after the dependency that caused the failure is healthy.
  • Roll back the application or Route configuration when a known-good revision exists.

Do not delete and recreate the Domain, certificate, or Route as a first response. That destroys relationship history and may introduce a second outage.

Test from an external client, then confirm the canonical hostname, certificate chain, status code, body, redirects, WebSockets, health history, and nginx logs. Close the incident only when public behavior and Opfield state agree.

Rollback should reverse the smallest change that introduced the outage. For DNS, restore the recorded address values and account for resolver caches. For TLS, restore the previous valid certificate assignment without disabling HTTPS. For nginx configuration, return to the last known-good managed settings and wait for acknowledgement. For an application regression, roll back the owning Deployment or Compose revision rather than rewriting the Route around a broken release.

Secure Link failures require restoring Relay and Node connectivity or the target runtime; changing DNS or certificates cannot repair them. Access-policy failures require restoring the previous Access List or trusted-proxy interpretation, not opening the workload port. Maintenance mode is useful while repairing these layers because it keeps the hostname and TLS path explicit while preventing accidental traffic to a partially recovered application.

When the first recovery attempt does not explain the failure, collect a bounded evidence bundle before escalating:

  • Domain and Route identifiers, assigned Ingress Node, and latest acknowledged config revision;
  • external DNS answers and TLS certificate details observed at the incident time;
  • Route health history and the exact failing path, method, status, and timestamp;
  • relevant nginx access/error entries and the owning workload or Secure Link operation;
  • Node and Relay versions, connectivity state, and the last failed Task message;
  • the last known-good application, Route, and DNS revision.

Redact credentials, cookies, authorization headers, private keys, and sensitive query values. The goal is to preserve causality without turning an incident record into a secret store.