Hserver Failure Notes: Safe Automation · intermediate

Read-Only Posture Checks Are Useful Because They Are Safe to Run Often

Observability improves when operators can refresh evidence without opening a risky change window.

Current. Current engineering note based on recent hserver deployment, debugging, recovery, and production-hardening work in September 2026.

Backup, DR, endpoint and production checks were most useful when operators could run them whenever evidence became stale. If collecting evidence changed the system, every diagnostic action would carry additional risk. Observation and mutation had not always been treated as separate operational capabilities. Safe diagnostics need different privilege and approval semantics from repairs. The portal exposes fixed LOW-risk posture jobs that read state, calculate evidence and write only their execution records.

Keep observation jobs side-effect free, document their evidence sources and create separate approved repair jobs when mutation is required. SRE troubleshooting works best when gathering evidence is cheap and repeatable. Read-only diagnostics reduce the temptation to 'fix while looking' before the failure has been localized. The concrete hserver evidence is commit 3c31a91, so this note is tied to an actual production change rather than a hypothetical failure.

Quick navigationEsc