Skip to main content

Retell vs. Vapi vs. LiveKit for voice AI agents: A comparison for developers

Let’s be honest, there isn’t a one-size-fits-all platform for developing voice agents. And to make matters a bit more blurry, if you’re an engineer comparing Retell, Vapi, and LiveKit, each is going to easily handle the standard voice AI demo.

We believe the differences between each platform begin to show themselves after the demo: when someone goes to iterate on the agent or once the agent is deployed at production scale.

Instead of doing a giant feature comparison between each, we’ve attempted to simplify the comparison into four key questions for engineering teams:

  1. Will engineers keep iterating on the agent after launch, or does it need to be iterated and updated without them?
  2. What happens when a requirement or optimization falls outside the defaults?
  3. Will the agent need to work beyond phone calls, on web, mobile, video, devices or WhatsApp?
  4. Is cost at scale a relevant concern?

Quick comparison#

We’ve included a high-level table below to outline some of the differences between Retell, Vapi, and LiveKit, but we suggest you dive into the details to understand the nuance.

Consideration factorRetellVapiLiveKit
Product categoryManaged voice-agent platformDeveloper-oriented managed voice platformDeveloper-oriented open source framework and cloud platform
Who owns the agent after launchOps or business teams, in a visual editor; developers via an API that mirrors itA developer, through config, APIs and webhooks; routine tuning in the dashboardEngineers who keep iterating on the agent as code
Primary interface for buildingDrag-and-drop workflow editorAssistant config via dashboard or API, plus webhooksPython and Node.js agent SDKs, a CLI to deploy. Includes an Agent Builder for getting started
ChannelsPhone (numbers included), SMS, web chat widgetPhone (free inbound number), web SDK, chat APIPhone (numbers included), web, mobile, video, devices and robots, WhatsApp, multi-party rooms
Open source and self-hostingNoNoYes, Apache 2.0; run on LiveKit Cloud or your own infrastructure
Platform costs as you scale*Flat per-minute rate; roughly 3.5x LiveKit at 10k minutes and aboveFlat per-minute rate; roughly 3.5x LiveKit at 10k minutes and aboveSimilar to the others at 1k minutes; about a quarter of their cost at 10k on Cloud, and no platform fee if you self-host

* Platform, telephony and observability fees on public non-enterprise list prices, September 2026, with your own SIP provider. Excludes model costs and carrier charges, which are roughly similar on all three. Skip to question 4 for more detailed numbers and assumptions.

Four questions to help you evaluate Retell vs. Vapi vs. LiveKit#

1. Will engineers keep iterating on it after launch, or does it need to run without them?#

Short answer

Retell is built so a non-engineer can improve the agent, which works until the change isn't a field Retell exposes. Vapi lets a developer script changes while Vapi runs the call, which means developers live inside Vapi's schema and versioning. LiveKit’s open source Agent framework puts the agent in your repo, which gives maximum flexibility for iterations and optimizations, but assumes you have engineers to keep iterating on it.

The reality is that, even for seemingly straightforward tasks, there's no such thing as the perfect AI voice agent out of the box. Any use beyond the demo will surface things the demo didn't such as:

  • A caller who pauses mid-sentence and gets cut off
  • A tool that takes three seconds at p95 and leaves dead air
  • A voicemail that gets a full sales pitch
  • A transfer to a human that drops the conversation context.

If you want to optimize your agent to achieve its goal, every agent you build with any platform will inevitably need to be tweaked in month two. The question is who makes them and what a change looks like.

RetellVapiLiveKit
The agent isConfiguration (dashboard first; API exposes the same fields)Configuration (API or dashboard)A program you deploy
A change isEdit a draft, diff, publish, roll out by %Edit a draft, publish, split traffic by % (beta), or PATCH the assistantPull request, lk agent deploy, rollback if wrong
Who builds itOps or marketingUsually a developerA developer
TestingText-only simulationsVoice + chat simulationsFull-pipeline audio simulations
TradeoffMore flexibility for non-developers, but less flexibility for developersSome flexibility for developer configuration, some flexibility for non-developersFull flexibility for developers, but it always requires one

2. What happens when a requirement or optimization falls outside the defaults?#

Short answer

The difference comes in how developers make a change and their limits. On Retell, you stop where the slider or field stops. On Vapi, you PATCH the assistant until the schema runs out. On LiveKit, it's a code change to the agent itself in your repo. The hosted platforms ship more out of the box, and if your workflow fits perfectly within their presets you might launch faster. However, many agents outgrow presets, and that's usually where working in code pays off. Instead of hunting for a setting that might do what you need, you write the change, test it like the rest of your code, and ship it using your git-based workflows (often with a coding agent doing much of the work).

Defaults and packaged integrations help teams get standard workflows running quickly. In some cases, though, they come at the cost of developer experience: the engineers responsible for improving the agent end up working inside a “black box.”

For basic scheduling, FAQ and lead-qualification agents, defaults are often the right trade, and both Retell and Vapi add packaged helpers like Retell's managed knowledge base and eleven first-party connectors for CRMs, calendars, helpdesks and knowledge sources or Vapi's “Campaigns” feature. The question is what happens when a requirement sits half a step outside the default workflow. In those instances, developers (and their coding agents) reap the benefits of a fully flexible open source framework like LiveKit’s.

Below is a non-comprehensive list meant to illustrate examples of what we mean by “configuration” and “defaults” versus “flexibility”.

I want to…RetellVapiLiveKit
Control how fast the agent replies and when it can be interruptedWait-time and interruption sliders, per agent or per flow node; built-in end-of-turn detection you can't swapChoose the turn model, set wait times by punctuation or regex, and tune interruption by word count and secondsFull flexibility to choose how turns are detected (with the built-in AI model, silence, or transcript). You can also tune min/max wait and interruption thresholds.
Remove background noise and other speakers from the caller's audioOne setting, three options: off, noise only (default), or noise + background speech (+metered)Krisp denoising (noise and background voices) on or off, plus an experimental Fourier filter you tune by dBChoose noise suppression (free) or voice isolation (metered), Krisp or ai-coustics, a telephony-tuned variant for SIP, adjust suppression level mid-call, or pick a different model per participant
Redact or edit what the caller said before the model sees itPost-call only, or host your own LLM serverDeepgram redaction at transcription, or host your own LLM serverA hook in your process, no network hop
Keep the conversation going while a long-running tool runsAgent fills the wait and answers if the caller speaks; the result lands in the transcript but isn't read back. Fields on the tool, set in the dashboard or agent JSONAgent holds the turn. Fields on the tool JSON, set in the dashboard or via APIYou decide, in the tool function itself (Python or Node). Agent can keep talking and speak the result when it lands. Full flexibility to choose how you handle async tools
Let a human screen the transferA screening agent talks with the human, then bridges or cancels. A decline or timeout returns the caller to your agent, without the human's reason. Dashboard or agent JSONHuman hears a message or summary, then is bridged. Talking back or declining needs the experimental mode. Configured in transferPlan JSONHuman can question the agent, connect, decline with a reason, or flag voicemail. Each outcome returns to your code. One awaited task, extend with your own tools

Compare the code: five real-world scenarios#

The table above outlines some examples of defaults vs. flexibility of the platforms, but developers know push comes to shove when you need to actually implement, read, and iterate the code. To help developers understand in more detail what it’s like to work with each platform, we’ve outlined a few of the scenarios below with some example code.

Scenario 1: A caller pauses mid-sentence and the agent cuts them off

"My account number is… hang on…" and the agent is already answering a question the caller hadn't finished. Every platform's defaults are tuned for a caller who talks in one steady stream. Fixing it for the caller who doesn't is usually the first optimization a team makes, and a good test of how far each platform's knobs reach.

Retell gives you a Response Wait time slider (0 to 5.5 seconds; responsiveness, 0 to 1, in the API), interruption_sensitivity, and an enable_dynamic_responsiveness toggle that adapts to the caller's pace. Retell already waits longer when it thinks the caller hasn't finished, but you can't choose or tune that detection. Switch STT to custom mode and you can also set the provider's endpointing_ms. To wait longer only after the agent asks for a number, build a Conversation Flow and override those settings on the node that collects it.

1
// PATCH https://api.retellai.com/update-agent/{agent_id}
2
{
3
"responsiveness": 0.6,
4
"interruption_sensitivity": 0.7,
5
"enable_dynamic_responsiveness": true,
6
"stt_mode": "custom",
7
"custom_stt_config": { "provider": "deepgram", "endpointing_ms": 800 }
8
}
9
// Agent-level: applies to every turn.
10
// A Conversation Flow node can override responsiveness and stt_mode for one step.

Vapi exposes far more in startSpeakingPlan: pick a smart endpointing provider (Vapi's docs recommend LiveKit's turn detector for English) or, without one, set separate wait times for turns that end on punctuation, no punctuation, or a number, and add regex rules that override either so the assistant waits longer after it says "account number." stopSpeakingPlan tunes interruption by word count and seconds. It's a PATCH to the assistant or a dashboard edit, and for this scenario it's enough. The ceiling is that the rule has to be expressible as a regex and a number of seconds.

1
// PATCH https://api.vapi.ai/assistant/{assistant_id}
2
{
3
"startSpeakingPlan": {
4
"waitSeconds": 0.4,
5
"smartEndpointingPlan": { "provider": "livekit", "waitFunction": "200 + 8000 * x" },
6
"customEndpointingRules": [
7
{
8
"type": "assistant",
9
"regex": "account number",
10
"regexOptions": [{ "type": "ignore-case", "enabled": true }],
11
"timeoutSeconds": 3
12
}
13
]
14
},
15
"stopSpeakingPlan": { "numWords": 2, "voiceSeconds": 0.3, "backoffSeconds": 1 }
16
}

LiveKit treats turn-taking as a pipeline stage you configure in code: choose the detector (the turn-detector model, VAD, STT endpointing, or manual), set min_delay and max_delay, switch endpointing to dynamic so the delay tracks the caller's actual pauses, and use adaptive interruption so an "uh-huh" doesn't stop the agent. Because it's code, the rule can be anything your app knows. Here the agent that collects the number changes the timing on entry, and it takes effect on the caller's next turn:

1
from livekit.agents import Agent, AgentSession, TurnHandlingOptions, inference
2
3
session = AgentSession(
4
# ... stt, llm, tts
5
turn_handling=TurnHandlingOptions(
6
turn_detection=inference.TurnDetector(), # or "vad", "stt", or "manual"
7
endpointing={"mode": "dynamic", "min_delay": 0.5, "max_delay": 3.0},
8
interruption={"mode": "adaptive"}, # an "uh-huh" doesn't stop the agent
9
),
10
)
11
12
# Handed off from your main agent when it's time to collect the number
13
class AccountNumberAgent(Agent):
14
def __init__(self) -> None:
15
super().__init__(
16
instructions="Ask for the account number, then read it back to confirm."
17
)
18
19
async def on_enter(self) -> None:
20
# Callers pause mid-number. Wait longer before ending their turn.
21
# Takes effect on the caller's next turn.
22
self.session.update_options(endpointing_opts={"min_delay": 1.5, "max_delay": 5.0})
23
await self.session.generate_reply(instructions="Ask for the account number.")
24
25
async def on_exit(self) -> None:
26
self.session.update_options(endpointing_opts={"min_delay": 0.5, "max_delay": 3.0})
Scenario 2: A caller reads out an order number and the transcript comes back garbled

The caller says "K7 dash 3R9Q." The transcript says "k7 3 are 9 queue." What happens next is where the platforms split.

Retell hands the transcript to the model. You can boost keywords in the STT, which helps, but the model still has to guess the order number, call your lookup tool with its guess, and ask the caller to repeat when the lookup fails. Fixing the transcript before the model sees it means taking over generation with a custom LLM server.

Vapi works the same way. Keyword boosting on Deepgram improves recognition, but the model still receives whatever the STT produced and does the guessing. Transforming the turn before the model means hosting a custom transcriber or a custom LLM server that Vapi streams through.

LiveKit edits the turn inside the agent process. You already know the caller's phone number, so you know their open orders. Fuzzy-match the garbled transcript against that list, then replace the turn with the result: "Order K7-3R9Q, shipped Tuesday, arriving Thursday." The model never guesses, never repeats a digit string, and the caller never hears "can you say that again?" No network hop, no second server.

1
import difflib
2
import re
3
4
from livekit.agents import Agent, ChatContext, ChatMessage
5
6
def normalize(text: str) -> str:
7
return re.sub(r"[^A-Z0-9]", "", text.upper())
8
9
class OrderStatusAgent(Agent):
10
async def on_user_turn_completed(
11
self, turn_ctx: ChatContext, new_message: ChatMessage
12
) -> None:
13
# Runs after STT and before the LLM, in the agent process. No network hop.
14
orders = self.session.userdata.open_orders # looked up by caller ID at call start
15
spoken = normalize(new_message.text_content or "")
16
candidates = {normalize(o.id): o for o in orders}
17
match = difflib.get_close_matches(spoken, list(candidates), n=1, cutoff=0.6)
18
if not spoken or not match:
19
return # nothing order-like; the model handles the turn as-is
20
21
order = candidates[match[0]]
22
# Replace the turn. The model relays a verified status instead of guessing digits.
23
new_message.content = [
24
f"What's the status of order {order.id}? "
25
f"[Verified from the caller's account: {order.status}, arriving {order.eta}]"
26
]
Scenario 3: Handle whatever picks up an outbound call: a person, voicemail, a phone menu, or nothing

Leaving a message is one field on every platform. Telling a menu from a mailbox, getting through the menu, and logging why 30% of last night's dials didn't connect is where they split.

Retell runs two detectors. Voicemail can hang up or leave a message; the IVR detector's only action is hang up, and navigating the menu is a separate feature you turn on. The two are handled separately, so a call classified as IVR never triggers the voicemail action

1
// PATCH https://api.retellai.com/update-agent/{agent_id}
2
{
3
"voicemail_option": {
4
"action": {
5
"type": "prompt",
6
"text": "Leave a short message for {{name}} asking them to call us back at {{callback_number}}."
7
}
8
},
9
"ivr_option": { "action": { "type": "hangup" } }
10
}
11
12
// To navigate a menu instead of hanging up, leave ivr_option unset and give the
13
// Retell LLM a press_digit tool. PATCH https://api.retellai.com/update-retell-llm/{llm_id}
14
{
15
"general_tools": [{ "type": "press_digit", "name": "press_digit", "delay_ms": 1000 }]
16
}
17
18
// The call webhook reports one reason per dial:
19
// "disconnection_reason": "voicemail_reached" | "ivr_reached" | ...
20
// A call classified as IVR never triggers the voicemail action.

Vapi picks a detection provider and returns one verdict, voicemail or not. However, a menu, a full mailbox and silence all read as "not voicemail," so the agent starts its pitch. Getting through a menu depends on the model deciding to call the dtmf tool.

1
// PATCH https://api.vapi.ai/assistant/{assistant_id}
2
{
3
"voicemailDetection": {
4
"provider": "vapi",
5
"backoffPlan": { "startAtSeconds": 2, "frequencySeconds": 2.5, "maxRetries": 5 },
6
"beepMaxAwaitSeconds": 20
7
},
8
"voicemailMessage": "Hi, this is Acme calling for {{name}}. Please call us back at 555-0100.",
9
"model": {
10
"provider": "openai",
11
"model": "gpt-4.1",
12
"messages": [
13
{
14
"role": "system",
15
"content": "... If you hear a phone menu, stay silent until the options finish, then use the dtmf tool (for example keys \"w2\") to reach billing."
16
}
17
],
18
"tools": [{ "type": "dtmf" }]
19
}
20
}
21
22
// Then POST https://api.vapi.ai/call with phoneNumberId, customer.number and assistantId.
23
// The end-of-call report gives one verdict: "endedReason": "voicemail", or not.
24
// A menu, a full mailbox and silence all read as "not voicemail", so the assistant
25
// starts talking, and getting through the menu depends on the model calling dtmf.

LiveKit classifies with an STT and LLM you pick and a prompt you can edit, returns one of five verdicts (person, voicemail, IVR menu, unavailable line, or uncertain) with the transcript behind it, and in Python starts IVR navigation on its own. Each verdict is a branch in your code and ops can get five reasons a dial failed, not one.

1
import os
2
3
from livekit import api
4
from livekit.agents import AMD, AgentSession, JobContext
5
6
SIP_TRUNK_ID = os.environ["LIVEKIT_SIP_OUTBOUND_TRUNK"]
7
8
async def dial(ctx: JobContext, session: AgentSession, phone_number: str) -> None:
9
identity = "callee"
10
session.room_io.set_participant(identity)
11
12
# Start AMD before the callee joins so no early audio is lost. The agent's
13
# speech is paused until the verdict is in.
14
async with AMD(session, participant_identity=identity) as detector:
15
await ctx.api.sip.create_sip_participant(
16
api.CreateSIPParticipantRequest(
17
room_name=ctx.room.name,
18
sip_trunk_id=SIP_TRUNK_ID,
19
sip_call_to=phone_number,
20
participant_identity=identity,
21
wait_until_answered=True,
22
)
23
)
24
await ctx.wait_for_participant(identity=identity)
25
26
result = await detector.execute() # STT and LLM you pick, prompt you can edit
27
log_dial_outcome(phone_number, result.category, result.transcript) # your ops report
28
29
if result.category in ("human", "uncertain"):
30
pass # the paused reply plays and the conversation continues
31
elif result.category == "machine-ivr":
32
pass # IVR navigation already started (ivr_detection=True is the default)
33
elif result.category == "machine-vm":
34
handle = session.generate_reply(
35
instructions="Leave a brief message asking the customer to call back."
36
)
37
await handle.wait_for_playout()
38
ctx.shutdown("voicemail")
39
elif result.category == "machine-unavailable":
40
ctx.shutdown("mailbox full or not set up") # retry later
Scenario 4: Let a human screen the transfer

All three can brief the human before connecting. The difference is what happens after the human hears the briefing.

Retell has two modes. Warm transfer speaks a one-shot whisper, then bridges; the human can't reply. Agentic warm transfer hands off to a separate screening agent you build, which talks with the human and then bridges or cancels. A cancel or timeout sends the caller back to your agent, but your agent never learns what the human said.

1
// On the Retell LLM. PATCH https://api.retellai.com/update-retell-llm/{llm_id}
2
{
3
"general_tools": [
4
{
5
"type": "transfer_call",
6
"name": "transfer_to_specialist",
7
"description": "Transfer the caller to a billing specialist after they confirm.",
8
"transfer_destination": { "type": "predefined", "number": "+14155550100" },
9
"transfer_option": {
10
"type": "warm_transfer",
11
"agent_detection_timeout_ms": 30000,
12
"on_hold_music": "relaxing_sound",
13
"show_transferee_as_caller": true,
14
"private_handoff_option": {
15
"type": "prompt",
16
"prompt": "Whisper a one-sentence summary of the caller's issue to the specialist."
17
},
18
"public_handoff_option": {
19
"type": "static_message",
20
"message": "You're now connected."
21
}
22
}
23
}
24
]
25
}
26
// Warm transfer: the whisper is one-way. Agentic warm transfer adds a screening agent.
27
// If nobody picks up within agent_detection_timeout_ms, the tool fails and
28
// your agent carries on with the caller.

Vapi has eight transfer modes. Seven are one-way: two blind, four that speak a fixed message or a generated summary, and one that runs your TwiML on the specialist's leg, then bridge. For a human who can ask a question or decline, you switch to the eighth, warm-transfer-experimental, and prompt a second assistant to run the handoff; if the human declines or doesn't answer, the caller comes back to your assistant or the call ends, per your fallback plan.

1
// In the assistant's model.tools, or POST https://api.vapi.ai/tool
2
{
3
"type": "transferCall",
4
"destinations": [
5
{
6
"type": "number",
7
"number": "+14155550100",
8
"description": "Transfer to a billing specialist after the caller agrees",
9
"transferPlan": {
10
"mode": "warm-transfer-experimental",
11
"transferAssistant": {
12
"firstMessage": "Hi, I have a caller with a billing question on hold. Can you take it?",
13
"firstMessageMode": "assistant-speaks-first",
14
"maxDurationSeconds": 120,
15
"silenceTimeoutSeconds": 30,
16
"model": {
17
"provider": "openai",
18
"model": "gpt-4o",
19
"messages": [
20
{
21
"role": "system",
22
"content": "Brief the specialist and answer their questions. Call transferSuccessful when they accept. Call transferCancel on voicemail, no answer, or a decline."
23
}
24
]
25
}
26
},
27
"holdAudioUrl": "https://example.com/hold.mp3",
28
"fallbackPlan": {
29
"message": "I couldn't reach a specialist. I can keep helping you.",
30
"endCallEnabled": false
31
}
32
}
33
}
34
]
35
}
36
// The other seven modes are one-way. This one runs a second assistant that
37
// talks to the specialist; on decline or no answer the caller gets the
38
// fallbackPlan message and either returns to your assistant or the call ends.

LiveKit runs the handoff as one awaited task. The human lands in a private room with the transcript and three tools: connect to the caller, decline with a reason, or flag voicemail. Whatever they choose comes back to your code.

1
import os
2
3
from livekit.agents import Agent
4
from livekit.agents.beta.workflows import WarmTransferTask, WorkflowInstructions
5
from livekit.agents.llm import ToolError, function_tool
6
7
SPECIALIST_NUMBER = os.environ["SPECIALIST_PHONE_NUMBER"]
8
9
class SupportAgent(Agent):
10
@function_tool
11
async def transfer_to_specialist(self) -> None:
12
"""Transfer the caller to a billing specialist. Only call after the caller confirms."""
13
await self.session.say(
14
"Please hold while I check if a specialist is available.",
15
allow_interruptions=False,
16
)
17
try:
18
# One awaited task: dials the specialist into a private room, plays hold music
19
# to the caller, briefs the specialist from chat_ctx, and gives them three tools:
20
# connect_to_caller, decline_transfer(reason), voicemail_detected.
21
await WarmTransferTask(
22
sip_call_to=SPECIALIST_NUMBER,
23
sip_trunk_id=os.environ["LIVEKIT_SIP_OUTBOUND_TRUNK"],
24
chat_ctx=self.chat_ctx,
25
ringing_timeout=25.0,
26
instructions=WorkflowInstructions(
27
extra="Say who is calling and what they need, then answer questions."
28
),
29
)
30
except ToolError as e:
31
# No answer, "human agent declined to connect: <reason>", or "voicemail detected".
32
# The caller is already off hold. The reason goes back to the model as the tool result.
33
raise ToolError(f"Could not transfer: {e}. Apologize and offer to take a message.") from e
34
35
await self.session.say(
36
"You're connected with a specialist now. Goodbye.", allow_interruptions=False
37
)
38
self.session.shutdown()
Scenario 5: Dial a list of patients tomorrow morning and report who picked up

This is the one where the hosted platforms often offer a better out-of-the-box experience.

Retell and Vapi ship it as a product. Retell's Batch Call takes a list with per-recipient variables, a start time, a call window and reserved concurrency, and shows Sent, Picked Up and Successful in the dashboard. Vapi's Campaigns take up to 10,000 contacts, a schedule window, a concurrency cap, per-contact webhooks and a pre-dial eligibility hook. Neither needs code.

As an open source framework, LiveKit Agents has no campaign product. Each dial is one dispatch against a SIP trunk you bring, and the list, the schedule, the pacing and the report are yours. It’s a real afternoon of work that the hosted platforms save you. What you get for it is no ceiling: pacing, retries and reporting are whatever you write, and concurrency is a plan tier rather than a per-line fee.

3. Will the agent need to work beyond phone calls, on web, mobile, video, devices or WhatsApp?#

Short answer

All three answer a phone. Retell adds SMS, MMS and a web chat widget as products. Vapi adds SMS through your Twilio number and a chat API. LiveKit is the only one where the model can see video (like a screen-share), a call can have more than two participants, the agent can run on hardware that isn't a phone or a browser, or answer a WhatsApp call. If the roadmap ends at phone plus text, any of the three works. If it includes web, mobile, video or devices, this question decides it.

ChannelRetellVapiLiveKit
Phone numbersBuy US or Canada numbers, inbound and outbound, calls to 15 countries. SIP import for the restOne free US number, inbound only. Outbound means importing a Twilio, Vonage, Telnyx or DIDWW number, or a SIP trunkOne free US local number per plan, inbound only today. Outbound and other countries run on a SIP trunk you configure
SMS and MMSA product on Retell or Twilio numbers. Two-way, MMS in, agent can text first or mid-callGlue over your Twilio account. Customer-initiated only, US to US, 24-hour sessions, no MMS. A built-in tool texts mid-callSending via Twilio or others is a function tool. Receiving is a small webhook you host that feeds a text-only session. Any carrier, agent can text first, MMS in as an image the model sees
Web and in-app chat (typed, not SMS)Chat-and-voice widget, callback widget, chat APIWidget (voice or text mode, one at a time) and chat APIVoice-first widget with text pane. Text-only widget: build it from the starter
Web and mobileBrowser SDK only, currentWeb SDK current. Flutter and React Native last shipped mid-2025; iOS has no tagged releaseTen client SDKs, all pushed regularly
Video in (the model sees)Images, audio or video texted in mid-call (MMS), which the agent can describeNo. Screen share and recording reach the call, not the modelFrames into any vision LLM; live camera or screen share with Gemini Live and OpenAI Realtime (Python)
AvatarsNoneTavus (though no longer in the docs)16 plugins
Devices and roboticsNonePython client for desktopESP32, embedded Linux, Jetson; ROS and LeRobot through LiveKit Portal
WhatsAppVia automation platformsVia automation platformsVoice calls through the WhatsApp Connector (Cloud)
More than two participantsOnly through transferOnly through transferRooms hold any number

The table is the quick version. Here’s each channel in more detail, and what it takes to support it on each platform.

Which platforms include phone numbers?#

All three, with different limits.

Retell is the only one that sells outbound-ready numbers with no telephony account. LiveKit and Vapi each include a free US number, and both are inbound-only.

Outbound on either means bringing a carrier. Vapi imports the number and handles the rest. LiveKit exposes the SIP layer: you configure trunks and dispatch rules, pick carrier and transport, and dial with CreateSIPParticipant. LiveKit is also the only one that answers WhatsApp voice calls natively, through a Cloud-only connector.

Can the agent send and receive SMS?#

All three, in different ways: Retell as a product, Vapi through Twilio, and LiveKit with a small amount of code.

Retell is the only one with SMS as a product: two-way texting on its own numbers, MMS attachments the agent can read, agent-initiated outbound, and A2P registration handled in the dashboard.

Vapi's SMS is a thin layer over a Twilio account you bring. It sets the webhook, keeps a 24-hour session per customer, and replies. It's Twilio-only, US-only, customer-initiated only, and has no MMS.

LiveKit doesn't ship SMS as a separate product, but adding it takes little code. Sending a text is a function tool that calls your carrier's API, and receiving one is a short webhook that passes the message into a text-only agent session. Because it's your code, it works with any carrier, the agent can text first or mid-call, and an MMS photo goes to the model as an image.

Can customers chat with the agent by typing on a website?#

All three, and each widget now handles voice too.

Retell's widget does text chat and browser voice calls, and Vapi's runs in either voice or text mode, one at a time. Both add a chat API. LiveKit's hosted widget is voice-first with a text pane. A text-only widget is two flags on the agent and a fork of the open-source embed starter. It’s buildable, but not prebuilt for you.

Which platforms have mobile SDKs, and are they maintained?#

All three run in a browser; only LiveKit's mobile SDKs are current.

Retell's client SDK is browser-only, though actively maintained. Vapi lists Web, iOS, Flutter, React Native and Python. The web package is current, but (as of October 2026) Flutter and React Native were last published in mid-2025, the iOS SDK has no tagged release, and the Python client hasn't shipped since 2024.

LiveKit ships ten client SDKs (web, Swift, Android, Flutter, React Native, Unity, Unity WebGL, C++, Rust, ESP32), each usually updated weekly.

Can the agent see video or images?#

Only on LiveKit.

In an STT-LLM-TTS pipeline you sample frames from the user's camera or screen into any vision-capable LLM. With Gemini Live or OpenAI Realtime, live video streams to the model directly (Python). Vapi's web SDK can start a screen share and record the call, but nothing actually reaches the model. Retell has no video plane. Its one media path is MMS during a phone call: images, audio or video the caller texts in, which the agent can describe.

Which platforms support video avatars?#

LiveKit has 16 avatar plugins.

Vapi's Tavus integration has dropped out of its documentation and survives only as credential and voice-provider types in the API. Retell has none.

Can the agent run on hardware, including a robot?#

Only LiveKit.

It runs on ESP32, embedded Linux and Nvidia Jetson, so the SDKs behind a phone agent also power a speaker, a kiosk or a wearable. For robots, LiveKit streams camera, sensor and control data in the same room as the voice agent, and LiveKit Portal syncs frames with robot state, arbitrates control between operator and policy, and bridges ROS and LeRobot.

Can more than two people be on the call?#

Only on LiveKit.

A room holds any number of humans and agents, which is how warm transfer, supervisor barge-in, and multi-agent calls work. On Retell and Vapi a call is one caller and one assistant; a second human only arrives through a transfer, and there's no documented way to add a third party and keep the agent on the line.

4. Is cost at scale a relevant concern?#

Short answer

Around 1,000 minutes a month, cost alone shouldn’t decide this; all three land in a similar range on platform costs. Above it, they split, and Retell and Vapi run 3.5–4x LiveKit's cost (platform, telephony and observability fees; excludes model costs and carrier charges, which are similar on all three). LiveKit also gives you two levers the others don't: cut the LLM bill to a fraction of a cent with fast open models like Gemma 4 through LiveKit Inference, or self-host the whole stack and pay no platform fee at all.

At the end of the day, most teams want a measurable ROI from their investment in voice. In many cases, the biggest return comes from two things: whether your agent delivers a successful experience to your end users, and whether the platform makes it easier for your developers to build, iterate, deploy, and keep optimizing it, which saves engineering time. At scale, though, cost starts to weigh more heavily, and it’s where the platforms diverge most.

Every voice agent bill has the same four parts: a platform fee for running the agent, model costs for STT, LLM and TTS (or a speech-to-speech model), telephony, and whatever you pay for logs and recordings. The model and carrier costs are roughly the same wherever you run them. The platform fee is where the vendors differ, and it's the part that scales with your minutes.

What you pay the platform#

Below are public list prices as of September 2026. Each assumes you bring your own SIP provider (Twilio, Telnyx, etc.). Totals only sum public plan fees, per-minute session fees, platform SIP fees and observability fees. Excludes model costs (STT, LLM, TTS; Retell bundles STT in its session fee), carrier charges, additional concurrency and enterprise discounts, which are similar or negotiable on all three.

Monthly minutesRetellVapiLiveKit Cloud
1,000$55$50$50 (Ship)
10,000$550$500$145 (Ship)
100,000$5,500$5,000$1,400 (Scale)

Why the gap opens up#

Retell and Vapi charge a flat per-minute platform fee. It's the same rate at your first thousand minutes and your millionth, so the bill grows in a straight line with usage.

LiveKit Cloud charges a lower per-minute rate plus a plan fee, and each plan includes a block of minutes before per-minute charges start. At low volume the plan fee makes LiveKit cost about the same as the others. As usage grows, the included minutes and the lower rate pull the per-minute cost down, which is why the gap widens with every order of magnitude. Past the point where any vendor's list price applies, all three negotiate, and the more relevant LiveKit option becomes self-hosting, where the platform fee is zero.

Model costs: same providers, different billing#

Bring the same STT, LLM and TTS to each platform and the per-minute cost is similar at launch. It diverges as the agent matures, because prompts grow: more instructions, more tools, a knowledge base.

Retell prices each model per minute and scales the bill up once your prompt passes a token threshold, counting instructions, tool definitions, the transcript so far and knowledge base content. Vapi passes through the provider's per-token price. LiveKit Inference bills per token, charges repeated prompt context at the provider's cached rate, and serves fast open-weight models like Gemma 4 for a fraction of a cent per minute. Switching models is one line of code.

Concurrency#

All three include a concurrency allowance and charge for more. Retell and Vapi sell additional capacity per line per month, and with burst turned on, Retell adds a $0.10/min surcharge to calls over your limit. LiveKit's allowance scales with plan tier and can be raised on Scale, with no per-line fee.

Compliance#

All three support HIPAA workloads, but at different price points: Retell includes a BAA on pay-as-you-go, LiveKit includes HIPAA and region pinning from the Scale plan up, and Vapi sells HIPAA mode as a monthly add-on. Vapi and LiveKit offer EU regions; Retell doesn't publish one.

Self-hosting#

Only LiveKit is open source. Run it yourself and there's no platform fee and no concurrency quota from the vendor; you pay for your own infrastructure and any Cloud-only pieces you keep, like Inference, phone numbers and observability. Retell and Vapi have no self-hosted option.

Bottom line: If volume is low and will stay low, the first three questions decide it. If you're heading past 1,000 or 10,000 minutes a month, expect your prompts to grow, or need the audio path in your own infrastructure, LiveKit's pricing is built for that, and it's the only one of the three you can run without the vendor.

Putting it together: which platform fits which team?#

There's no wrong answer here, only a wrong fit. Map your team against the four questions and one platform usually falls out. For most teams, the deciding one is the second: how often you'll need to go past the defaults.

If your team…Start with
Is building a product where voice is the experience: language learning, coaching, tutoring, companionship, gaming, a deviceLiveKit
Has no engineers on the agent after launch, and ops or marketing will own and iterate itRetell
Has a developer who can configure and wire up webhooks, the workflow fits a standard pattern, and volume is moderate to lowVapi
Has engineers who will keep iterating, and needs to tune, extend or scale beyond the defaultsLiveKit
Needs web, mobile, video, devices or more than two people on a callLiveKit
Expects heavy usage, or needs the audio path in its own infrastructureLiveKit

Choose Retell if the agent needs to run without engineers. It's the only one of the three where the whole agent lives in a visual editor a non-developer can own, and it ships the most as products: numbers, SMS, batch calling, a knowledge base. The ceiling is whatever the editor exposes.

Choose Vapi if a developer prefers to assemble the agent from configuration, the defaults are good enough, and the use case follows a well-worn path: simple scheduling, qualification, campaigns. Vapi keeps adding rails for those patterns. You're opting into them rather than writing and optimizing the behavior for your own use case.

Choose LiveKit if your developers would rather build the agent in code than in configuration. With LiveKit, the agent is a program in your repo, so changes go through pull requests, code review, tests and CI like the rest of your stack, and coding agents like Claude Code, Cursor and Codex can build and change it directly. That matters most once you go past the defaults (which most production agents do). It's also the stronger fit if cost at scale matters. The tradeoff is that there's no visual editor or pre-packaged outbound calling campaign product.

One more cut that's easy to miss: Retell and Vapi are built for customer-facing calls. Their feature sets (transfers, voicemail detection, campaigns, SMS follow-up, batch dialing) are primarily contact-center features, and the unit of work is focused around a phone call with one caller and one assistant. If you're building a language-learning app, a voice coaching product, a tutoring or companionship experience, an in-game character, or a voice interface for a device, the product is the conversation itself and LiveKit is the only one of the three built to support that.

Frequently asked questions#

Which platform is fastest to set up: Retell, Vapi, or LiveKit?

All three can have a phone agent answering calls within a day. Retell is fastest for a team with no developer, since the whole agent is built in a visual editor. Vapi and LiveKit both take a developer an afternoon: Vapi via a config and a webhook, LiveKit via a ~30-line agent deployed to LiveKit Cloud. For an engineering team the difference is hours, not weeks; what separates the platforms is how quickly you can change the agent after launch.

Which platform gives you telephony out of the box?

All three. LiveKit is often assumed to require a carrier, but LiveKit Cloud provisions US local and toll-free numbers directly, includes a free local number on every plan, and answers inbound calls without a Twilio account. Outbound calls, transfers and international calls run over a SIP trunk from a provider you choose, and LiveKit is the only one of the three that answers WhatsApp voice calls natively. Retell sells outbound-ready numbers with no carrier account. Vapi's free numbers are inbound-only; outbound means importing a number from Twilio, Telnyx, Vonage or a SIP trunk.

How do you deploy, manage, and host LiveKit?

Run lk agent deploy from your project and LiveKit Cloud builds a container from your code, deploys it, and scales it with call volume. New versions roll out without dropping calls: new sessions go to the new build while active ones finish on the old one. Deployment logs, session metrics, recordings and traces are in the dashboard, with instant rollback and separate non-production deployments on paid plans. If you'd rather run the agent yourself, the same code deploys to Kubernetes, Render or any container platform, connected to LiveKit Cloud transport or a self-hosted LiveKit server.

Which platform has the lowest latency?

None of the three publishes a comparable end-to-end benchmark, and most of a voice agent's latency comes from the models: how fast the STT finalizes, how long the LLM takes to first sentence, how quickly the TTS starts speaking. All three let you pick fast models. Where they differ is what you can do about the rest. Retell exposes two sliders. Vapi exposes the turn model and a set of timing thresholds in config. LiveKit exposes every stage in code: how turns are detected, which turn model runs, endpointing and interruption thresholds, and where the agent runs.

Inference is the other lever. On LiveKit Cloud the agent is hosted next to the media server and LiveKit Inference, so model calls don't leave the cluster. And LiveKit hosts latency-optimized open-weight models, like Gemma 4 31B, on its own GPUs; on LiveKit's public live benchmark it returns a first sentence several times faster than the frontier APIs, and several times faster than the same model served through a general-purpose provider. The benchmark re-runs every ten minutes and shows p95 and p99 next to p50, because a fast median doesn't help the caller who hits the slow tail. If latency is a competitive feature of your product rather than a number to keep under a threshold, it's a difference that matters.

Can you use speech-to-speech models like OpenAI Realtime or Gemini Live?

Yes on all three, as a model choice. The difference is what you can do around it. On LiveKit, a realtime model is one option in the same framework as a pipeline agent, so you can pair it with a separate TTS voice, keep your own turn detection and noise cancellation in front of it, hand off between realtime and pipeline agents mid-call, and (in Python) stream live camera or screen-share video into it. On Retell and Vapi you select the realtime model in configuration and take the platform's handling of turns and audio.

Which platform should you use if voice is the product, not a support channel?

LiveKit. Retell and Vapi are designed around customer-service calls or outbound campaigns: a phone number, a caller, a resolution or a transfer. A language-learning app, a coaching product, a tutoring experience or a voice-driven game has different needs: it lives inside your web or mobile app, sessions run long, the model may need to see video, and there may be more than two participants. LiveKit's client SDKs, video support, multi-participant rooms and device targets are built for that, and the agent is code in your app, so the experience is yours to design.

Which is the most cost-effective at scale?

At a few thousand minutes a month, the three cost about the same on platform fees. Past roughly 10,000 minutes, Retell and Vapi's flat per-minute rates hold while LiveKit's included minutes and lower session rate pull its cost down, so LiveKit runs at a fraction of the others' platform cost by 100,000 minutes. Model costs are similar everywhere, with one exception: Retell scales its per-minute LLM price up as prompts grow, while LiveKit and Vapi bill per token. See question four for the numbers.

Can you self-host any of these?

Only LiveKit. The server, SIP service, egress and agents framework are open source under Apache 2.0, so you can run the whole stack on your own infrastructure and pay LiveKit nothing. What you give up is the Cloud-only layer: managed inference, observability, the dashboard, Agent Builder, Krisp noise cancellation and LiveKit's hosted turn-detection model. Most teams that self-host keep transport on LiveKit Cloud and run only the agent themselves, which removes the per-session fee without losing those features. Retell and Vapi have no self-hosted option.

Can a coding agent like Codex, Claude Code, or Cursor build on these?

All three publish an MCP server and machine-readable docs, so the answer is yes for each. The difference is what the coding agent is being asked to do. On Retell and Vapi, the MCP server lets an assistant drive the platform: create agents, adjust config, place calls. On LiveKit the coding agent writes the agent itself, and LiveKit publishes a docs MCP server, an llms.txt index, a CLI and agent skills that cover workflow design, handoffs and testing, so tools like Claude Code, Cursor and Codex can efficiently build, test, iterate, and deploy an agent end to end. If your plan is to have a coding agent own most of the implementation, the code-first platform is where it has the most to work with.

Related