Hserver Monitoring: Logs & Security · advanced

Kernel OOM Logs Are the Ground Truth for Host Memory Kills

A process disappearing can look like an application crash unless the kernel journal is checked for OOM-killer activity.

Current. Current production-engineering note derived from the hserver observability deployment, runtime measurements, alert rules, dashboards, and recovery work in September 2026.

A process disappearing can look like an application crash unless the kernel journal is checked for OOM-killer activity. On the finished hserver stack, kernel OOM event counters and matching journal records is the signal that makes the difference visible. Kernel logs identify when memory pressure caused the operating system to kill a process, separating resource failure from application logic failure.

The engineering pattern here is event-source correlation. Good monitoring should shorten diagnosis, so I prefer a small number of signals with clear semantics over a larger collection whose meaning is unclear during a failure.

Operationally I keep this constraint: Alert on OOM events immediately, link them with memory PSI and container OOM metrics, and preserve the surrounding journal context for postmortem analysis. It gives the dashboard, alert, and runbook the same interpretation instead of letting each layer invent its own definition of healthy.

The implementation can be traced to hserver commit b65d5d4. That provenance is part of the article because these notes document an actual production observability system, not a hypothetical monitoring design.

Quick navigationEsc