Perspectives · 8 min read · 2025-06-19
A scrub detects latent corruption, a resilver reconstructs missing redundancy, a snapshot preserves an earlier dataset state, and replication puts that state somewhere else. Calling all four 'backup' hides the failures each one cannot solve.
Disaster Recovery · 1 min read · 2025-06-08
A failed job can leave a fresh-looking directory that should never replace the last known complete recovery point.
Production Engineering · 1 min read · 2024-08-16
Live files that are absent from Git are technical debt even when the service is currently healthy.
Disaster Recovery · 1 min read · 2024-05-06
Separating secrets from source control creates a recovery dependency that must be documented and tested.
LOUP Engineering · 1 min read · 2024-03-25
Production fixtures need a recovery path below the application firmware.
Terraform Systems · 8 min read · 2023-07-21
State maps resource addresses in configuration to real remote objects. Treating it as disposable cache data is how infrastructure gets orphaned or recreated.
Terraform Systems · 8 min read · 2023-04-04
Resource targeting narrows Terraform to a selected subset plus dependencies. It is valuable for exceptional recovery but can leave the rest of the configuration unapplied.
Security & Identity · 1 min read · 2022-12-22
Losing the authenticator device creates a second-factor recovery problem that should not silently collapse to the password path.
Terraform Systems · 8 min read · 2022-12-03
state mv, state rm, state replace-provider, state pull and state push can change Terraform's ownership model without directly changing remote infrastructure.
Production Engineering · 1 min read · 2021-12-30
A scheduler can enqueue work before crashing; recovery logic must discover incomplete deliveries after the process returns.
LOUP Engineering · 1 min read · 2020-08-12
A device that loses power or Wi-Fi halfway through setup should not become permanently ambiguous.