2025 Voice and Infrastructure Notes · advanced

The ESP32-S3 Voice Pipeline Was Really Several Clocks Sharing One Device

Microphone sampling, RTP packetization, network arrival and speaker playout each have their own timing domain.

Lab Note. Reconstructed in 2026 as a technical lab note from the 2025 voice, telecom and infrastructure study period. The archive date indicates the period covered; the current site publication date is shown separately.

A voice pipeline on the ESP32-S3 looked linear on paper: microphone to codec, codec to RTP, RTP over Wi-Fi, then the reverse path for playback. The bugs made more sense when I stopped treating it as one timeline.

The microphone samples from a hardware clock. RTP packets are grouped on a packetization interval. The network adds variable arrival time. The speaker consumes audio on another hardware clock. If those rates differ slightly, buffers drift even when there is no packet loss.

That explains why a call can sound good for thirty seconds and then slowly begin to underrun or accumulate latency. The fix is not always a larger buffer. I need one part of the system to own playout timing and explicit policies for drift, rebuffering and late packets.

Once I viewed the pipeline as clock domains connected by buffers, audio debugging became much less random. The traces started matching what I could hear instead of leaving me to guess between network and codec problems.

Quick navigationEsc