Partial Backup Directories Should Not Win the 'Newest Backup' Contest
A failed job can leave a fresh-looking directory that should never replace the last known complete recovery point.
TOPIC
Lessons, explainers, experiments, and implementation notes.
A failed job can leave a fresh-looking directory that should never replace the last known complete recovery point.
Scheduling evidence and completion evidence belong to different layers of a batch job.
When opening a file is slow, the disk is an obvious suspect. But data may cross a filesystem, mount, network gateway, and application layer before reaching the user. The symptom does not prove that storage is where the delay began.
A recovery drill can create a new sensitive-data exposure if decrypted database dumps remain on disk by default.
Separating secrets from source control creates a recovery dependency that must be documented and tested.
A backup on the same disk protects against application mistakes better than it protects against host or disk loss.
Recovery planning becomes actionable when data-loss tolerance and recovery-time tolerance are explicit.
Protected evidence that a low-privilege checker cannot read should not be reported as healthy or corrupt.
Knowing that a backup file exists matters, but whether the required system can be rebuilt from it is a restore question. Having a copy and having a working recovery path are two different guarantees.
A newly created directory can still be incomplete or corrupt, so age alone is weak recovery evidence.
Checksums do not replace restore tests, but they catch silent byte changes before a disaster forces you to discover them.
The backup process can exit zero while the recovery process is still incomplete, undocumented or impossible on another machine.
Stopping only the shell does not necessarily stop pg_dump, tar or helper processes that the shell launched.
A backup that hangs forever can block future runs and create false confidence without ever producing a usable recovery point.
Rollback usually sounds like restoring an older release. But if the new code changed the structure or meaning of data, whether the old code can still read that data is a separate question.
Operators need evidence that backups succeeded without necessarily gaining access to the protected data itself.