Hserver Monitoring: Alerting & Notification · advanced
Alert Delivery Failure Needs an Alert of Its Own
A monitoring system that detects failures but cannot notify anyone is partially failed even if every Prometheus target remains green.
A monitoring system that detects failures but cannot notify anyone is partially failed even if every Prometheus target remains green. On hserver the first signal I use for this question is HserverExternalNotificationDeliveryFailed from alert-sink failure counters. The notification path is a production dependency and must be monitored from inside the observability system rather than assumed to work forever.
The important part is interpretation rather than collecting another graph. meta-monitoring of paging infrastructure. That gives the metric a specific operational job instead of making it another number on a dashboard.
The practical control is straightforward: Page or surface notification failures through an alternate visible channel and regularly test delivery with controlled synthetic alerts. This also gives me a repeatable check after deployments, exporter changes, or capacity tuning.
Repository evidence for this monitoring behavior is commit 49ec1dd. I keep that reference with the note because a monitoring conclusion is stronger when the configuration and runtime decision that produced it can be inspected later.