Hserver Monitoring: Alerting & Notification · advanced

Alert Delivery Failure Needs an Alert of Its Own

A monitoring system that detects failures but cannot notify anyone is partially failed even if every Prometheus target remains green.

Current. Current production-engineering note derived from the hserver observability deployment, runtime measurements, alert rules, dashboards, and recovery work in September 2026.

A monitoring system that detects failures but cannot notify anyone is partially failed even if every Prometheus target remains green. On hserver the first signal I use for this question is HserverExternalNotificationDeliveryFailed from alert-sink failure counters. The notification path is a production dependency and must be monitored from inside the observability system rather than assumed to work forever.

The important part is interpretation rather than collecting another graph. meta-monitoring of paging infrastructure. That gives the metric a specific operational job instead of making it another number on a dashboard.

The practical control is straightforward: Page or surface notification failures through an alternate visible channel and regularly test delivery with controlled synthetic alerts. This also gives me a repeatable check after deployments, exporter changes, or capacity tuning.

Repository evidence for this monitoring behavior is commit 49ec1dd. I keep that reference with the note because a monitoring conclusion is stronger when the configuration and runtime decision that produced it can be inspected later.

Quick navigationEsc