Block I/O per Container Exposes Noisy Neighbors
Host disk latency can rise because one container is performing heavy reads or writes while every other service only sees the consequence.
Host disk latency can rise because one container is performing heavy reads or writes while every other service only sees the consequence.
Old Docker volumes can consume storage after services are removed, and their names alone do not always reveal whether anything still depends on them.
Swap is sometimes treated as a direct replacement for RAM, but memory moved to storage has to be read back when active work needs it again. Swap creates room while introducing a time cost.
Logs are useful for future investigations, but unlimited retention can let the evidence system consume the resources required by the service itself. Observability needs its own storage budget.
A scrub detects latent corruption, a resilver reconstructs missing redundancy, a snapshot preserves an earlier dataset state, and replication puts that state somewhere else. Calling all four 'backup' hides the failures each one cannot solve.
A disk can have plenty of free capacity while the underlying SSD is reporting temperature or device-health problems.
A database normally grows, so alerting on size alone would create noise while ignoring the important question of growth rate and disk headroom.
It is easy to think of saving data as one indivisible event: write it and it is done. In reality, power can disappear in the middle of a write. The next boot then has to distinguish complete, incomplete, and usable state.
When opening a file is slow, the disk is an obvious suspect. But data may cross a filesystem, mount, network gateway, and application layer before reaching the user. The symptom does not prove that storage is where the delay began.
A database doing many tiny synchronous operations and a backup streaming large files can show similar disk utilization with very different access patterns.
Communication is simpler if every node is assumed to be continuously available. Real nodes disappear because of power, radio conditions, or the environment. Then the design has to decide where a message waits and how long it remains relevant.
Keeping thirty days of Prometheus data is useful until series growth causes the TSDB to consume more disk than the host can safely spare.