Skip to main content

Benchmarks

Voice AI, measured from every angle. From latency to cost to real-world capabilities, see exactly how each model performs.

Latency over time

02505007501,0001 AM4 AM7 AM10 AM1 PM4 PM7 PM10 PM
Live

Full results

Creator
Model name
Service tier
Gemma
Gemma 4 31B
235 ms594 ms1,174 ms247 ms669 ms1,478 ms317 ms753 ms1,517 ms262 tok/s97 tok/s34 tok/s0.4%4,940
Google
Gemini 2.5 Flash
priority773 ms1,009 ms1,554 ms765 ms1,093 ms1,557 ms814 ms1,113 ms1,654 ms219 tok/s118 tok/s99 tok/s0.1%4,940
Gemini 2.5 Flash
fast610 ms1,598 ms2,678 ms610 ms1,592 ms2,635 ms664 ms1,660 ms2,804 ms227 tok/s130 tok/s71 tok/s0.0%4,940
OpenAI
GPT-4.1
fast526 ms806 ms1,133 ms635 ms1,058 ms1,437 ms731 ms1,167 ms1,608 ms145 tok/s76 tok/s47 tok/s0.0%4,940
GPT-4.1
priority472 ms946 ms2,206 ms604 ms1,222 ms2,380 ms657 ms1,303 ms2,678 ms177 tok/s78 tok/s48 tok/s0.0%4,940

Stream is the generation rate in tokens/sec of decode time.
Higher is better. P5 + P1 are the worst-case figures that determine whether the model stays ahead of speech playback.

Measured from LiveKit infrastructure, so latencies include the network path to each provider.
They are intended for relative comparison, not as an absolute SLA.

methodology

Start simulating

Measurement

We record time to first token, time to first complete sentence, total response time, and generation throughput. First sentence matters most for voice: it's when the agent can start talking, which is what a caller experiences as response delay, and it can sit well behind first-token latency on a model that opens with a short fragment. We measure network round-trip separately, so provider latency reads apart from distance, and we count failures and timeouts rather than dropping them, making error rate part of the result instead of a gap. Turns are aggregated hourly and reported as medians and 95th/99th percentiles.

Conversational Replay

Every 10 minutes, from a single vantage point, we replay the same scripted customer-support conversation against every model and provider. The script is sized like a real deployment: a ~3,500-token system prompt with policies, FAQs and examples, a nine-message support call, and 5 tool definitions. 5 user turns are measured as the conversation grows, and two require the model to pick the right tool out of the five. The numbers reflect both a lengthening prompt and the tool-calling work a real agent does.

Controlled Comparison

Every target sees byte-identical requests: the same prompt, the same tools, and the same generation settings (128 max tokens, temperature 0.7, top-p 1.0), sent over the same public HTTPS endpoints any customer would use. The benchmark refuses to start if any of that drifts apart between targets.

Benchmark your own agents

Spin up simulated conversations, run repeatable benchmarks, and measure exactly how your agents perform under real-world conditions.