Skip to content

Diagnose unexpected monitoring health

Start with the selected account, check kind, latest result timestamp and individual location/resolver results. Aggregate health is a rollup, not a complete diagnosis.

Symptom Check first Corrective action
Heartbeat never leaves pending Correct location URL and an actual accepted check-in Send a success-only report from the intended sender
Heartbeat becomes late Job logs, next deadline, cron timezone and grace Fix job failure or align expectations with real completion
Healthy timing but degraded/down Numeric values and quorum Compare inclusive warning/critical thresholds and every location
HTTP failure Expected code, redirects, public target, TLS and WAF Inspect the recorded response/error; test the exact URL
TCP failure Hostname, public port, listener and firewall Confirm a public connection is allowed
DNS mismatch Resolver answers, type, full expected set and mode Correct expectations; inspect TTL and lagging state
Expiry date missing/stale Latest lookup and actual date availability Verify hostname and registrar/TLS endpoint; do not infer expiration from a lookup error
Health unchanged after a network error Infrastructure-suspect result and freshness Look for a subsequent trustworthy result

HTTP 200 from ingestion is not proof the job succeeded. exit_code is recorded, not a status failure instruction. A malformed numeric value is ignored without rejecting the check-in and does not erase the last value. Shared location URLs allow one sender to hide another’s silence.

Website checks do not render JavaScript or check response-body text. Private targets and private redirect destinations are blocked. DNS includes compares normalized records, not substrings. A resolver serving previous expected records during an applicable TTL lag window can remain healthy.

Do not disable alerts to hide a genuine monitoring error. Fix the target, expectation or routing independently and verify a subsequent result.