2021 SIP Infrastructure · intermediate

My First SIPp Load Test Was Mostly About Defining a Useful Scenario

SIPp could generate a lot of calls, but the hard part was deciding what behavior to simulate and what failure actually meant.

Lab Note. Reconstructed in 2026 as a lab note from the 2021 VoIP and SIP-infrastructure study period. The archive date indicates the period covered; the current site publication date is shown separately.

SIPp made it easy to create impressive-looking call rates, which was exactly why I had to be careful with it. A load test is only useful when the scenario resembles the behavior I want to measure. Sending thousands of synthetic INVITEs without realistic timing, responses or media says very little about a production call path.

I started with a small UAC/UAS scenario. The caller sent INVITE, handled provisional and final responses, sent ACK, waited for a defined call duration, then sent BYE. The answering side responded predictably. Before increasing the rate, I verified one call in a packet trace so I knew the scenario itself was correct.

Then I increased calls per second gradually and watched more than CPU. I tracked response times, failed calls, retransmissions, socket errors and transaction timeouts. A proxy can show moderate CPU while a downstream database, PBX or media component is already becoming the bottleneck. Successful process uptime is not the same thing as successful call handling.

Media changes the test substantially. A signaling-only scenario measures signaling capacity. If the real system also anchors RTP, transcodes codecs or records calls, that resource cost has to be tested separately or included deliberately. Otherwise the result is a capacity number for a different architecture.

I also learned to distinguish call rate from concurrency. Ten calls per second with two-second duration creates a different steady-state load from ten calls per second with five-minute duration. Registration storms are another workload entirely. Naming the workload precisely made later performance discussions much more useful.

The first failures were valuable. Some came from file-descriptor limits and local test-generator constraints rather than the SIP server. That reminded me to monitor both the system under test and the load generator. A benchmark is meaningless if the generator cannot produce the requested traffic cleanly.

SIPp became most useful when I stopped asking 'how many calls can this server handle?' and started asking narrower questions: at this call rate and call duration, with this routing path and these media assumptions, when do latency and failure rate become unacceptable? That is a test I can reproduce and compare after changing the system.

Quick navigationEsc