Clocked Playout Beats Packet-Driven Audio
Network packets arrive when the network allows; speakers need samples when the audio clock demands them.
Network packets arrive when the network allows; speakers need samples when the audio clock demands them.
Containerizing a SIP service added another address and NAT boundary, which made port publishing, advertised addresses and RTP ranges more important rather than less important.
A small closed call loop isolates core SIP/RTP behavior before carrier and DID complexity is introduced.
Two softphones and one Asterisk server were enough to show that a phone call is really several network problems stacked together: registration, signaling, media and NAT.
Every extra buffered frame trades conversational responsiveness for tolerance to arrival variation.
Media ports must agree across PBX configuration, firewall policy, containers and the surrounding network.
Media sessions can close for rejection, timeout, silent timeout, final timeout or offer timeout, and those reasons point to different failure paths.
A running SIP or RTP process is necessary but not sufficient evidence that calls can establish and carry media.
A long-running packet tool looked suspicious, but removing it did not fully restore 20 ms scheduling.
A larger jitter buffer can hide network variation while quietly making conversation worse through extra delay.
Kamailio could route the signaling perfectly while media still failed. rtpengine made the signaling path and media path explicit instead of treating them as one thing.
REGISTER proves one control-plane exchange; it does not prove two-way media, codecs or call-state behavior.
Network arrival time is not an audio clock. A stable playout schedule needs its own timing and buffer policy.
A real endpoint pair exposes assumptions hidden by softphones and local echo tests.
A secure SIP transport and encrypted media are separate decisions. TLS can protect signaling while RTP remains completely visible on the wire.
Engineering time was being spent chasing tens of milliseconds in firmware while the media route crossed Bangladesh and Ohio twice.
Microphone sampling, RTP packetization, network arrival and speaker playout each have their own timing domain.
A low average loss rate can still sound terrible when packets arrive in short bursts separated by long gaps.
A SIP call can connect perfectly and still carry no useful audio if the negotiated media description is wrong.
Burst timing can hurt a speakerphone even when aggregate RTP loss is close to zero.
FreeSWITCH forced me to separate the concepts I understood from Asterisk from the implementation details I had simply memorized.
After server cleanup, Asterisk forwarding became fast, but conversation still felt delayed.
Media cannot be debugged reliably if the firewall and PBX disagree on where RTP is allowed.
Signaling and media usually take different paths, use different ports and fail for different reasons, so a working SIP registration says very little about RTP health.
Packet timing only becomes actionable when it is tied to endpoint tolerance.
The device received media in repeating bursts that looked like a local queue or I2S starvation problem.
Audio packets do not arrive at perfectly regular intervals even when they travel the same network path. A jitter buffer waits long enough to smooth that variation, but the waiting itself increases conversational latency.
Short audible failures can be network gaps, scheduling gaps or local conversion artifacts.
A signalling-only ladder can say a call succeeded while the user heard silence; adding SDP and RTP events fixes that blind spot.
Anchoring media at a deliberate boundary reduced the number of private addresses that leaked into SDP and simplified firewall policy.
Following one VoLTE call from LTE attachment through IMS registration, SIP session setup, bearer creation and RTP finally connected the year's separate labs.
A timestamp inside a packet can look like a clock reading, but RTP timestamps primarily describe position in the media clock. Wall-clock time, media time, and packet arrival time are different concepts.
If one side hears audio, several parts of the media path are already proven.
More buffering reduces underruns until it starts creating latency and hiding drift.
Server scheduling can inject media timing problems even when embedded firmware has not changed.
RTP arrival time should not directly schedule the speaker.
When the call timer starts, we tend to assume communication has succeeded. But SIP can establish the signaling state while the path that carries audio is still broken. Who is talking to whom and how the media packets reach them are not answered by the same process.
A SIP call can signal perfectly and still have broken audio because SDP is where the endpoints describe media addresses, ports and codecs.
A packet can arrive at 09:17:25 and carry an RTP timestamp like 2873419200. Those values describe different things. Confusing them is one of the fastest ways to make real-time audio timing harder than it already is.
One-way audio was my first VoIP problem where the signaling looked healthy and the real fault was the address and port information used for RTP.
Packet loss, reordering and burst arrival can sound similar but require different fixes.
Once SIP, RTP, codecs, Wi-Fi and a UI share one MCU, memory and timing decisions stop being implementation details.
On a headless PBX, a narrow tcpdump capture often answered the important question faster than a full GUI trace: did the signaling or media packet actually reach the server?
Asterisk was configured for 20 ms media timing, yet packet forwarding arrived in scheduler-sized bursts.