Hserver Monitoring: Host & Resource Signals · advanced

Critical systemd Units Need Their Own Health Signal

Container monitoring does not cover host services such as Docker, networking, tunnels, backup timers or other systemd-managed dependencies.

Current. Current production-engineering note derived from the hserver observability deployment, runtime measurements, alert rules, dashboards, and recovery work in September 2026.

Container monitoring does not cover host services such as Docker, networking, tunnels, backup timers or other systemd-managed dependencies. What made the issue measurable was hserver_systemd_unit_active and hserver_systemd_unit_failed. A healthy container stack can still be unusable when a required host unit is inactive, failed or repeatedly restarting.

I classify this as dependency-aware service monitoring. The useful debugging sequence is to confirm the signal, compare it with the neighboring subsystem, then look at logs or detailed metrics only after the failure domain is smaller.

The production rule that came out of it is: Maintain a reviewed allow-list of critical host units and alert on state changes instead of scraping every systemd unit indiscriminately. This is deliberately more specific than adding another broad alert with no response procedure.

Commit b65d5d4 is the repository evidence behind the note. It provides the concrete configuration or fix that turned the observation into a repeatable monitoring control.

Quick navigationEsc