Benchmarks
Voice AI, measured from every angle. From latency to cost to real-world capabilities, see exactly how each model performs.
Latency over time
Creator | Model name | Service tier | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Gemma 4 31B | — | 235 ms | 594 ms | 1,174 ms | 247 ms | 669 ms | 1,478 ms | 317 ms | 753 ms | 1,517 ms | 262 tok/s | 97 tok/s | 34 tok/s | 0.4% | 4,940 | |
Gemini 2.5 Flash | priority | 773 ms | 1,009 ms | 1,554 ms | 765 ms | 1,093 ms | 1,557 ms | 814 ms | 1,113 ms | 1,654 ms | 219 tok/s | 118 tok/s | 99 tok/s | 0.1% | 4,940 | |
Gemini 2.5 Flash | fast | 610 ms | 1,598 ms | 2,678 ms | 610 ms | 1,592 ms | 2,635 ms | 664 ms | 1,660 ms | 2,804 ms | 227 tok/s | 130 tok/s | 71 tok/s | 0.0% | 4,940 | |
GPT-4.1 | fast | 526 ms | 806 ms | 1,133 ms | 635 ms | 1,058 ms | 1,437 ms | 731 ms | 1,167 ms | 1,608 ms | 145 tok/s | 76 tok/s | 47 tok/s | 0.0% | 4,940 | |
GPT-4.1 | priority | 472 ms | 946 ms | 2,206 ms | 604 ms | 1,222 ms | 2,380 ms | 657 ms | 1,303 ms | 2,678 ms | 177 tok/s | 78 tok/s | 48 tok/s | 0.0% | 4,940 |
Stream is the generation rate in tokens/sec of decode time.
Higher is better. P5 + P1 are the worst-case figures that determine whether the model stays ahead of speech playback.
Measured from LiveKit infrastructure, so latencies include the network path to each provider.
They are intended for relative comparison, not as an absolute SLA.
Start simulating
Measurement
We record time to first token, time to first complete sentence, total response time, and generation throughput. First sentence matters most for voice: it's when the agent can start talking, which is what a caller experiences as response delay, and it can sit well behind first-token latency on a model that opens with a short fragment. We measure network round-trip separately, so provider latency reads apart from distance, and we count failures and timeouts rather than dropping them, making error rate part of the result instead of a gap. Turns are aggregated hourly and reported as medians and 95th/99th percentiles.
Conversational Replay
Every 10 minutes, from a single vantage point, we replay the same scripted customer-support conversation against every model and provider. The script is sized like a real deployment: a ~3,500-token system prompt with policies, FAQs and examples, a nine-message support call, and 5 tool definitions. 5 user turns are measured as the conversation grows, and two require the model to pick the right tool out of the five. The numbers reflect both a lengthening prompt and the tool-calling work a real agent does.
Controlled Comparison
Every target sees byte-identical requests: the same prompt, the same tools, and the same generation settings (128 max tokens, temperature 0.7, top-p 1.0), sent over the same public HTTPS endpoints any customer would use. The benchmark refuses to start if any of that drifts apart between targets.
Benchmark your own agents
Spin up simulated conversations, run repeatable benchmarks, and measure exactly how your agents perform under real-world conditions.