2026 Production Voice Systems · advanced

RTP Playout Got Better When I Stopped Letting Packet Arrival Drive the Speaker

Network arrival time is not an audio clock. A stable playout schedule needs its own timing and buffer policy.

Lab Note. Reconstructed in September 2026 from current engineering work and lab notes. The archive date indicates the period covered; the current site publication date is shown separately.

One of the worst ways to play RTP is also one of the easiest to implement: receive a packet and immediately push its samples toward the speaker. It sounds acceptable on a perfect LAN and falls apart as soon as packet arrival starts to bunch up.

The fix was to make playout clock-driven. RTP packets enter a jitter buffer, while a separate schedule consumes one audio frame at a time at the codec rate. Network arrival can then vary without directly changing the speaker timing.

That exposed the next problem: rebuffer policy. If the buffer drains, I need a deliberate threshold for pausing, refilling and resuming. If it grows continuously, there may be clock drift rather than ordinary jitter. Logging buffer depth over a long call was more useful than listening to a thirty-second test.

The important distinction is simple. The network transports audio frames; it should not be the clock that plays them. Giving playout its own timing source made the buffer measurable and turned random crackle into a scheduling problem I could actually instrument.

Quick navigationEsc