TOPIC

Observability & Monitoring

Lessons, explainers, experiments, and implementation notes.

Observability & Monitoring · 2 min read · 2026-08-01

Label Freedom Comes With a Time-Series Cost

Metric labels let us divide one measurement in many useful ways, but every new combination of label values can create another time series. Greater detail also carries collection, storage, and query cost.

Observability & Monitoring · 2 min read · 2025-12-15

Who Does What When an Alert Fires

The value of an alert is not limited to detecting failure. It also includes what someone is expected to do after the signal arrives. If ownership is vague, even an accurate alert becomes just another piece of noise.

Observability & Monitoring · 2 min read · 2025-11-16

When Clocks Drift, the Order of Events Can Drift Too

When logs from several servers are compared side by side, timestamps are often used to reconstruct the incident. If the clocks differ, however, a later event can appear to have happened first. Sorting timestamps is not enough to prove causal order.

Observability & Monitoring · 2 min read · 2025-06-24

Do Not Let Log History Block Current Work

Logs are useful for future investigations, but unlimited retention can let the evidence system consume the resources required by the service itself. Observability needs its own storage budget.

Observability & Monitoring · 2 min read · 2025-01-09

What Do We Lose When We Do Not Keep Every Trace?

Keeping a complete trace for every request can be expensive. Sampling controls that cost, but it also defines what an investigation can and cannot see. A missing trace is not evidence that the event never happened.

Observability & Monitoring · 2 min read · 2024-10-31

Which Identity Should Follow a Call Through the System?

A phone call that crosses several servers leaves pieces of its story in several places. Endpoint logs, PBX logs, and trunk logs may all describe it differently. Time proximity alone is not enough to prove that two log entries belong to the same call.

Observability & Monitoring · 2 min read · 2024-09-09

No Data Is Not Zero

Replacing missing data with zero can make a graph look cleaner while changing its meaning. Zero says a measurement was made and the result was zero. No data says the result is unknown.

Observability & Monitoring · 2 min read · 2024-04-15

The Slow Users Disappear Inside the Average

A low average response time can make a system look fast even when a smaller group of users waits much longer. Their experience disappears into one number. To understand latency, we also need to understand its distribution.

Observability & Monitoring · 2 min read · 2023-06-26

Who Monitors the Monitoring System?

If monitoring stops, evidence of failures elsewhere can disappear too. Silence can then look like health. The monitoring system therefore needs evidence that its own collection and notification paths are still working.

Observability & Monitoring · 2 min read · 2023-04-22

A Counter Reset Changes the Story in the Graph

A counter accumulates events over time, but some counters start over when a process restarts. Subtracting the next value from the previous one can then produce an impossible negative event. The events did not run backward; the measurement history was reset.

Quick navigationEsc