A Tinny Speaker Is Not Automatically a Codec Problem
Field feedback described the speaker as very tinny even though the digital call path was working.
TOPIC
Lessons, explainers, experiments, and implementation notes.
Field feedback described the speaker as very tinny even though the digital call path was working.
Network packets arrive when the network allows; speakers need samples when the audio clock demands them.
RTP arrived in bursts even when packet sequence was mostly healthy, so directly pacing I2S from receive callbacks made network jitter audible.
Prove clocks, framing, routing and gain before judging microphones, speakers or enclosure acoustics.
A processor that removes noise can also remove speech detail when its assumptions do not match the signal.
Every extra buffered frame trades conversational responsiveness for tolerance to arrival variation.
A long-running packet tool looked suspicious, but removing it did not fully restore 20 ms scheduling.
One strong call can prove a candidate is promising but not that it is ready for manufacturing or field release.
Echo cancellation only has a useful reference when the reference represents what the loudspeaker actually played.
A large first callback-to-speaker delay suggested work was accumulating between RTP reception and physical output.
A rebuffer/AEC/volume experiment produced obvious bad crackle instead of the intended stability improvement.
Later queue, PLC and timer experiments looked more sophisticated on paper but repeatedly introduced crackle, echo or additional delay.
A larger audio buffer can hide some problems temporarily, but extra capacity does not make shared memory safe when ownership is unclear. The design still needs rules for who writes, who reads, and when a region may be reused.
Engineering time was being spent chasing tens of milliseconds in firmware while the media route crossed Bangladesh and Ohio twice.
An I2S Mode1 conflict warning looked suspicious enough to become a candidate explanation for crackle.
It was tempting to use server packet timing as proof of the complete mouth-to-ear delay.
The AEC library consumed a block size that did not divide evenly into the telephony frame size.
Codec detection alone did not prove the capture channels meant what the DSP assumed they meant.
After digital timing became stable, echo increased at high speaker volume and could no longer be treated as only a network or queue problem.
After server cleanup, Asterisk forwarding became fast, but conversation still felt delayed.
The firmware sounded worse while producing the very diagnostics intended to explain it.
Logs help us observe a system, but producing them also consumes time. Heavy logging inside a real-time audio path can change the behavior being measured.
Words like robotic, delayed and crackly were useful user reports but poor root-cause evidence.
The device received media in repeating bursts that looked like a local queue or I2S starvation problem.
Audio packets do not arrive at perfectly regular intervals even when they travel the same network path. A jitter buffer waits long enough to smooth that variation, but the waiting itself increases conversational latency.
Crackle and cutouts needed a device-side timing metric that was closer to the DAC than packet arrival.
A call can sound excellent for thirty seconds while two media clocks slowly walk apart.
Voice-agent latency is often reduced to how quickly the model generates text. In reality, the conversation includes audio capture, network transport, end-of-speech detection, transcription, generation, speech synthesis, and playback.
Robotic speech, cutouts, lag and echo initially collapsed into one vague complaint called bad audio.
The physical capture clock and the network codec did not run at the same sample rate.
Fast-moving firmware work made it easy for an attractive theory to become remembered as fact.
Increasing jitter tolerance seemed like the obvious response to bursty packet arrival.
Bytes moving over I2S do not prove that both devices interpret those bytes the same way. Sample width, channel layout, clock relationships, and data alignment all have to match. A silent speaker is not automatically an amplifier problem.
The downlink could be packet-complete and still sound gritty or artificial after conversion to the physical playback rate.
Packet loss, reordering and burst arrival can sound similar but require different fixes.
A working phone is not a release unless the exact artifact and source context can survive the next experiment.
Asterisk was configured for 20 ms media timing, yet packet forwarding arrived in scheduler-sized bursts.
Echo tuning was meaningless if the AEC reference did not represent what the loudspeaker actually played.
AEC was implicated in several symptoms, but disabling it permanently would remove a required speakerphone function.
Later experiments had changed queues, PLC, timing and logs until nobody could safely say which behavior belonged to the last clear build.