Disk Usage Alerts Need Warning and Critical Bands
A single disk threshold gives operators no distinction between early cleanup work and a filesystem that is close to stopping writes.
A single disk threshold gives operators no distinction between early cleanup work and a filesystem that is close to stopping writes.
CPU and container counts are supporting signals; service reachability, backup validity and call-path health are closer to the real objective.
A production-engineering deep dive into thermals on a 2014 mac mini are a production signal, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
Container monitoring does not cover host services such as Docker, networking, tunnels, backup timers or other systemd-managed dependencies.
A production-engineering deep dive into linux psi changed how i think about saturation, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A panel can render old data successfully after the collector behind it has already stopped.
Label names and cardinality are query contracts for dashboards, recording rules and alerts.
Logs, TLS checks, backup ages and distributed event ordering all become harder to trust when the host clock drifts.
Observability can become the workload if collection is broader than the questions operators actually need to answer.
A production-engineering deep dive into disk throughput is not disk latency, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into recording rules are a cpu trade: compute once or query repeatedly, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into why every target should not be scraped every 15 seconds, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
Dashboards improved once I stopped collecting attractive metrics and started collecting evidence for specific failure modes.
A few megabytes of swap on an old Linux host did not automatically mean an incident, especially after long uptime.
A target that disappeared can leave its last sample available long enough for threshold expressions to evaluate against stale data.
A production-engineering deep dive into why i monitor memavailable instead of “free ram”, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
Not every metric needs to be collected at the same cadence; expensive inventory can be cached without weakening real-time health signals.
Labels make dashboards flexible, but uncontrolled label values multiply time series and memory cost quickly.
A production-engineering deep dive into 30-day retention and a 15 gb cap were capacity decisions, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
CPU can change meaningfully in seconds while Docker image storage or Grafana process memory does not need the same fifteen-second collection cadence.
A broken rule evaluator or missing target can produce a beautifully quiet alert dashboard while the system is blind.
A few retransmissions during heavy traffic may be normal, while the same count during low traffic can represent a serious quality problem.
A deployment may push a metric over threshold without immediately firing because the configured `for` period has not elapsed.
A load average of four means something very different on a two-core system than on an eight-core system.
The host looked busy enough that a simple percentage could easily become the whole diagnosis, but Linux memory reclaim makes that misleading.
The SIP proxy can be running as a process while its control connection to drachtio is down, leaving signaling logic unable to operate correctly.
A recent backup timestamp can look reassuring even when the archive is incomplete, corrupt or impossible to restore.
A service can look healthy now and still have restarted repeatedly overnight, erasing the evidence from a simple current-state view.
A single SIP proxy error may be harmless noise, but repeated failures over a short interval can indicate backend, routing or dependency trouble.
A production-engineering deep dive into load average without cpu count is almost meaningless, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
Caching Docker storage inventory reduced overhead, but a silent refresh failure could otherwise leave old values looking current indefinitely.
The hserver stack carried roughly twenty-seven thousand active Prometheus series, enough that label growth and exporter changes could materially change memory and storage cost.
A systemd timer can fire on schedule while the backup service itself fails, times out or exits before producing a valid set.
The edge process can be green while one proxied service is failing behind it.
A small Mac mini running many containers can hit thermal constraints before ordinary CPU graphs explain why performance changed.
Delivery counters only change when an alert is sent, so a quiet system could leave a dead notification service unnoticed for hours.
A production-engineering deep dive into clock synchronization is an availability dependency, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into monitoring tcp retransmission on a server that also runs voip, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into swap usage alone is a bad memory alert, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into rule evaluation failures mean monitoring logic is broken, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into prometheus is my metrics control plane, not just a scraper, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A count of sixty active PostgreSQL connections is meaningless until it is compared with the configured maximum for that server.
A production-engineering deep dive into 27.5k active series on a small server: what that number actually means, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into disk full and inodes full are two different outages, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into wal, head block and why prometheus memory does not equal stored data, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
CPU, packet loss and endpoint probes can cross thresholds for a few seconds during harmless transitions or deployment activity.
A production-engineering deep dive into cardinality is usually a schema problem, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
Logs explain individual failures; metrics show whether the system is drifting before the failures become obvious.
A production-engineering deep dive into metric relabeling is resource control, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
High interface traffic can be completely healthy, while a small but sustained drop rate can damage voice, APIs and tunnel reliability.
Grafana, Loki and Alloy expose many internal metrics that are useful for development but unnecessary on a small production Prometheus.
A PostgreSQL deadlock can resolve automatically by aborting one transaction, leaving the service apparently healthy after the incident.
A production-engineering deep dive into how much prometheus is too much for an 8 gb machine?, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
A production-engineering deep dive into page cache is not a memory leak, grounded in the 2014 Mac mini hserver observability stack and its accepted runtime evidence.
RX and TX graphs looked busy enough, but throughput by itself could not tell whether traffic was healthy.
Keeping thirty days of Prometheus data is useful until series growth causes the TSDB to consume more disk than the host can safely spare.