Skip to main content

Benchmarks

Voice AI, measured from every angle. From latency to cost to real-world capabilities, see exactly how each model performs.

Latency over time

Comparison of latency across top models. Lower scores are better.

Full results

methodology

Start simulating

Measurement

We record time to first token, time to first complete sentence, total response time, and generation throughput. First sentence matters most for voice: it's when the agent can start talking, which is what a caller experiences as response delay, and it can sit well behind first-token latency on a model that opens with a short fragment. We measure network round-trip separately, so provider latency reads apart from distance, and we count failures and timeouts rather than dropping them, making error rate part of the result instead of a gap. Turns are aggregated hourly and reported as medians and 95th/99th percentiles.

Conversational Replay

Every 10 minutes, from a single vantage point, we replay the same scripted customer-support conversation against every model and provider. The script is sized like a real deployment: a ~3,500-token system prompt with policies, FAQs and examples, a nine-message support call, and 5 tool definitions. 5 user turns are measured as the conversation grows, and two require the model to pick the right tool out of the five. The numbers reflect both a lengthening prompt and the tool-calling work a real agent does.

Controlled Comparison

Every target sees byte-identical requests: the same prompt, the same tools, and the same generation settings (128 max tokens, temperature 0.7, top-p 1.0), sent over the same public HTTPS endpoints any customer would use. The benchmark refuses to start if any of that drifts apart between targets.

Benchmark your own agents

Spin up simulated conversations, run repeatable benchmarks, and measure exactly how your agents perform under real-world conditions.