Measure What the Operator or User Actually Cares About
CPU and container counts are supporting signals; service reachability, backup validity and call-path health are closer to the real objective.
TOPIC
Lessons, explainers, experiments, and implementation notes.
CPU and container counts are supporting signals; service reachability, backup validity and call-path health are closer to the real objective.
A probe that follows a redirect can report success for the wrong endpoint and hide an authentication or routing failure.
Observability can become the workload if collection is broader than the questions operators actually need to answer.
Not every metric needs to be collected at the same cadence; expensive inventory can be cached without weakening real-time health signals.
Labels make dashboards flexible, but uncontrolled label values multiply time series and memory cost quickly.
The fastest observability optimization came from identifying one costly collector instead of globally lowering fidelity.
A last-known-good value becomes misleading when the system does not show how old the evidence is.
Hardware and drivers expose radio state differently, so a production metric may need a fallback without hiding uncertainty.