Hserver Failure Notes: Observability · intermediate

Stale Telemetry Should Not Render as Healthy

A last-known-good value becomes misleading when the system does not show how old the evidence is.

Current. Current engineering note based on recent hserver deployment, debugging, recovery, and production-hardening work in September 2026.

The Fleet and Command Center views could display reassuring state even when the runner or storage observation had not refreshed recently. The number itself was valid when collected, but the UI did not make evidence age prominent enough.

SRE monitoring depends on trustworthy evidence, not just positive values. Freshness is part of the signal contract, especially for safety gates and release decisions. State and freshness had been collapsed into one concept. A green health result from an hour ago is not equivalent to a green result from thirty seconds ago when the underlying service can change quickly.

Telemetry runners gained explicit states such as STALE and NEVER_SEEN, storage observations carry timestamps, and the operator UI flags stale data instead of presenting it as current truth.

Every derived health signal should have an age budget. If the evidence exceeds that budget, downgrade the state to stale or unknown and require a new observation before high-risk actions proceed. The concrete hserver evidence is commit be4f22b, so this note is tied to an actual production change rather than a hypothetical failure.

Quick navigationEsc