Hserver Monitoring: Docker & Containers · advanced

Container Restart Counts Are Incident Breadcrumbs

A service can look healthy now and still have restarted repeatedly overnight, erasing the evidence from a simple current-state view.

Current. Current production-engineering note derived from the hserver observability deployment, runtime measurements, alert rules, dashboards, and recovery work in September 2026.

A service can look healthy now and still have restarted repeatedly overnight, erasing the evidence from a simple current-state view. I ended up treating container restart counters and runtime state as the useful observation point rather than relying on a generic service-up indicator. Restarts are durable clues for crash loops, OOM kills, daemon restarts and unstable dependencies even after the process comes back.

This is a good example of event-history monitoring. The purpose is to reduce ambiguity during an incident: a signal should tell me which layer to inspect next, not simply confirm that something somewhere looks unusual.

For production I use the following guardrail: Graph restart deltas and correlate them with logs, OOM events and host pressure so recovered failures still trigger investigation. The same rule keeps the dashboard useful when the system grows and more targets are added.

The implementation is traceable to b65d5d4 in the hserver repository. That commit is the concrete reference for the collector, alert, dashboard, or runtime change behind this article.

Quick navigationEsc