Hserver Monitoring: Backup & DR · advanced
Backup Job Exit Status Belongs in Prometheus
A systemd timer can fire on schedule while the backup service itself fails, times out or exits before producing a valid set.
A systemd timer can fire on schedule while the backup service itself fails, times out or exits before producing a valid set. What made the issue measurable was hserver_systemd_service_result_success for managed backup units. Scheduling evidence and completion evidence are separate; monitoring only the timer would report success for a failed backup execution.
I classify this as job-outcome monitoring. The useful debugging sequence is to confirm the signal, compare it with the neighboring subsystem, then look at logs or detailed metrics only after the failure domain is smaller.
The production rule that came out of it is: Export the last service result, alert on failure, and correlate it with backup age so a failed run cannot remain hidden until the next day. This is deliberately more specific than adding another broad alert with no response procedure.
Commit 11ff139 is the repository evidence behind the note. It provides the concrete configuration or fix that turned the observation into a repeatable monitoring control.