Hserver Monitoring: VoIP · advanced

No Healthy SIP Worker Is a User-Facing Failure Counter

The most important load-balancer failure is not that a backend probe failed; it is that an incoming SIP request could not be assigned to any healthy worker.

Current. Current production-engineering note derived from the hserver observability deployment, runtime measurements, alert rules, dashboards, and recovery work in September 2026.

The most important load-balancer failure is not that a backend probe failed; it is that an incoming SIP request could not be assigned to any healthy worker. What made the issue measurable was voip_sip_proxy_no_healthy_worker_total. Counting rejected requests translates internal pool health into actual service impact and gives a stronger paging signal than backend state alone.

I classify this as impact-based alerting. The useful debugging sequence is to confirm the signal, compare it with the neighboring subsystem, then look at logs or detailed metrics only after the failure domain is smaller.

The production rule that came out of it is: Alert on any recent increase, then correlate with worker health and heartbeat age to identify whether capacity or routing caused the rejection. This is deliberately more specific than adding another broad alert with no response procedure.

Commit b65d5d4 is the repository evidence behind the note. It provides the concrete configuration or fix that turned the observation into a repeatable monitoring control.

Quick navigationEsc