Let’s be honest, there isn’t a one-size-fits-all platform for developing voice agents. And to make matters a bit more blurry, if you’re an engineer comparing Retell, Vapi, and LiveKit, each is going to easily handle the standard voice AI demo.
We believe the differences between each platform begin to show themselves after the demo: when someone goes to iterate on the agent or once the agent is deployed at production scale.
Instead of doing a giant feature comparison between each, we’ve attempted to simplify the comparison into four key questions for engineering teams:
- Will engineers keep iterating on the agent after launch, or does it need to be iterated and updated without them?
- What happens when a requirement or optimization falls outside the defaults?
- Will the agent need to work beyond phone calls, on web, mobile, video, devices or WhatsApp?
- Is cost at scale a relevant concern?
Quick comparison#
We’ve included a high-level table below to outline some of the differences between Retell, Vapi, and LiveKit, but we suggest you dive into the details to understand the nuance.
| Consideration factor | Retell | Vapi | LiveKit |
|---|---|---|---|
| Product category | Managed voice-agent platform | Developer-oriented managed voice platform | Developer-oriented open source framework and cloud platform |
| Who owns the agent after launch | Ops or business teams, in a visual editor; developers via an API that mirrors it | A developer, through config, APIs and webhooks; routine tuning in the dashboard | Engineers who keep iterating on the agent as code |
| Primary interface for building | Drag-and-drop workflow editor | Assistant config via dashboard or API, plus webhooks | Python and Node.js agent SDKs, a CLI to deploy. Includes an Agent Builder for getting started |
| Channels | Phone (numbers included), SMS, web chat widget | Phone (free inbound number), web SDK, chat API | Phone (numbers included), web, mobile, video, devices and robots, WhatsApp, multi-party rooms |
| Open source and self-hosting | No | No | Yes, Apache 2.0; run on LiveKit Cloud or your own infrastructure |
| Platform costs as you scale* | Flat per-minute rate; roughly 3.5x LiveKit at 10k minutes and above | Flat per-minute rate; roughly 3.5x LiveKit at 10k minutes and above | Similar to the others at 1k minutes; about a quarter of their cost at 10k on Cloud, and no platform fee if you self-host |
* Platform, telephony and observability fees on public non-enterprise list prices, September 2026, with your own SIP provider. Excludes model costs and carrier charges, which are roughly similar on all three. Skip to question 4 for more detailed numbers and assumptions.
Four questions to help you evaluate Retell vs. Vapi vs. LiveKit#
1. Will engineers keep iterating on it after launch, or does it need to run without them?#
Retell is built so a non-engineer can improve the agent, which works until the change isn't a field Retell exposes. Vapi lets a developer script changes while Vapi runs the call, which means developers live inside Vapi's schema and versioning. LiveKit’s open source Agent framework puts the agent in your repo, which gives maximum flexibility for iterations and optimizations, but assumes you have engineers to keep iterating on it.
The reality is that, even for seemingly straightforward tasks, there's no such thing as the perfect AI voice agent out of the box. Any use beyond the demo will surface things the demo didn't such as:
- A caller who pauses mid-sentence and gets cut off
- A tool that takes three seconds at p95 and leaves dead air
- A voicemail that gets a full sales pitch
- A transfer to a human that drops the conversation context.
If you want to optimize your agent to achieve its goal, every agent you build with any platform will inevitably need to be tweaked in month two. The question is who makes them and what a change looks like.
| Retell | Vapi | LiveKit | |
|---|---|---|---|
| The agent is | Configuration (dashboard first; API exposes the same fields) | Configuration (API or dashboard) | A program you deploy |
| A change is | Edit a draft, diff, publish, roll out by % | Edit a draft, publish, split traffic by % (beta), or PATCH the assistant | Pull request, lk agent deploy, rollback if wrong |
| Who builds it | Ops or marketing | Usually a developer | A developer |
| Testing | Text-only simulations | Voice + chat simulations | Full-pipeline audio simulations |
| Tradeoff | More flexibility for non-developers, but less flexibility for developers | Some flexibility for developer configuration, some flexibility for non-developers | Full flexibility for developers, but it always requires one |
2. What happens when a requirement or optimization falls outside the defaults?#
The difference comes in how developers make a change and their limits. On Retell, you stop where the slider or field stops. On Vapi, you PATCH the assistant until the schema runs out. On LiveKit, it's a code change to the agent itself in your repo. The hosted platforms ship more out of the box, and if your workflow fits perfectly within their presets you might launch faster. However, many agents outgrow presets, and that's usually where working in code pays off. Instead of hunting for a setting that might do what you need, you write the change, test it like the rest of your code, and ship it using your git-based workflows (often with a coding agent doing much of the work).
Defaults and packaged integrations help teams get standard workflows running quickly. In some cases, though, they come at the cost of developer experience: the engineers responsible for improving the agent end up working inside a “black box.”
For basic scheduling, FAQ and lead-qualification agents, defaults are often the right trade, and both Retell and Vapi add packaged helpers like Retell's managed knowledge base and eleven first-party connectors for CRMs, calendars, helpdesks and knowledge sources or Vapi's “Campaigns” feature. The question is what happens when a requirement sits half a step outside the default workflow. In those instances, developers (and their coding agents) reap the benefits of a fully flexible open source framework like LiveKit’s.
Below is a non-comprehensive list meant to illustrate examples of what we mean by “configuration” and “defaults” versus “flexibility”.
| I want to… | Retell | Vapi | LiveKit |
|---|---|---|---|
| Control how fast the agent replies and when it can be interrupted | Wait-time and interruption sliders, per agent or per flow node; built-in end-of-turn detection you can't swap | Choose the turn model, set wait times by punctuation or regex, and tune interruption by word count and seconds | Full flexibility to choose how turns are detected (with the built-in AI model, silence, or transcript). You can also tune min/max wait and interruption thresholds. |
| Remove background noise and other speakers from the caller's audio | One setting, three options: off, noise only (default), or noise + background speech (+metered) | Krisp denoising (noise and background voices) on or off, plus an experimental Fourier filter you tune by dB | Choose noise suppression (free) or voice isolation (metered), Krisp or ai-coustics, a telephony-tuned variant for SIP, adjust suppression level mid-call, or pick a different model per participant |
| Redact or edit what the caller said before the model sees it | Post-call only, or host your own LLM server | Deepgram redaction at transcription, or host your own LLM server | A hook in your process, no network hop |
| Keep the conversation going while a long-running tool runs | Agent fills the wait and answers if the caller speaks; the result lands in the transcript but isn't read back. Fields on the tool, set in the dashboard or agent JSON | Agent holds the turn. Fields on the tool JSON, set in the dashboard or via API | You decide, in the tool function itself (Python or Node). Agent can keep talking and speak the result when it lands. Full flexibility to choose how you handle async tools |
| Let a human screen the transfer | A screening agent talks with the human, then bridges or cancels. A decline or timeout returns the caller to your agent, without the human's reason. Dashboard or agent JSON | Human hears a message or summary, then is bridged. Talking back or declining needs the experimental mode. Configured in transferPlan JSON | Human can question the agent, connect, decline with a reason, or flag voicemail. Each outcome returns to your code. One awaited task, extend with your own tools |
Compare the code: five real-world scenarios#
The table above outlines some examples of defaults vs. flexibility of the platforms, but developers know push comes to shove when you need to actually implement, read, and iterate the code. To help developers understand in more detail what it’s like to work with each platform, we’ve outlined a few of the scenarios below with some example code.
Scenario 1: A caller pauses mid-sentence and the agent cuts them off
"My account number is… hang on…" and the agent is already answering a question the caller hadn't finished. Every platform's defaults are tuned for a caller who talks in one steady stream. Fixing it for the caller who doesn't is usually the first optimization a team makes, and a good test of how far each platform's knobs reach.
Retell gives you a Response Wait time slider (0 to 5.5 seconds; responsiveness, 0 to 1, in the API), interruption_sensitivity, and an enable_dynamic_responsiveness toggle that adapts to the caller's pace. Retell already waits longer when it thinks the caller hasn't finished, but you can't choose or tune that detection. Switch STT to custom mode and you can also set the provider's endpointing_ms. To wait longer only after the agent asks for a number, build a Conversation Flow and override those settings on the node that collects it.
1// PATCH https://api.retellai.com/update-agent/{agent_id}2{3"responsiveness": 0.6,4"interruption_sensitivity": 0.7,5"enable_dynamic_responsiveness": true,6"stt_mode": "custom",7"custom_stt_config": { "provider": "deepgram", "endpointing_ms": 800 }8}9// Agent-level: applies to every turn.10// A Conversation Flow node can override responsiveness and stt_mode for one step.
Vapi exposes far more in startSpeakingPlan: pick a smart endpointing provider (Vapi's docs recommend LiveKit's turn detector for English) or, without one, set separate wait times for turns that end on punctuation, no punctuation, or a number, and add regex rules that override either so the assistant waits longer after it says "account number." stopSpeakingPlan tunes interruption by word count and seconds. It's a PATCH to the assistant or a dashboard edit, and for this scenario it's enough. The ceiling is that the rule has to be expressible as a regex and a number of seconds.
1// PATCH https://api.vapi.ai/assistant/{assistant_id}2{3"startSpeakingPlan": {4"waitSeconds": 0.4,5"smartEndpointingPlan": { "provider": "livekit", "waitFunction": "200 + 8000 * x" },6"customEndpointingRules": [7{8"type": "assistant",9"regex": "account number",10"regexOptions": [{ "type": "ignore-case", "enabled": true }],11"timeoutSeconds": 312}13]14},15"stopSpeakingPlan": { "numWords": 2, "voiceSeconds": 0.3, "backoffSeconds": 1 }16}
LiveKit treats turn-taking as a pipeline stage you configure in code: choose the detector (the turn-detector model, VAD, STT endpointing, or manual), set min_delay and max_delay, switch endpointing to dynamic so the delay tracks the caller's actual pauses, and use adaptive interruption so an "uh-huh" doesn't stop the agent. Because it's code, the rule can be anything your app knows. Here the agent that collects the number changes the timing on entry, and it takes effect on the caller's next turn:
1from livekit.agents import Agent, AgentSession, TurnHandlingOptions, inference23session = AgentSession(4# ... stt, llm, tts5turn_handling=TurnHandlingOptions(6turn_detection=inference.TurnDetector(), # or "vad", "stt", or "manual"7endpointing={"mode": "dynamic", "min_delay": 0.5, "max_delay": 3.0},8interruption={"mode": "adaptive"}, # an "uh-huh" doesn't stop the agent9),10)1112# Handed off from your main agent when it's time to collect the number13class AccountNumberAgent(Agent):14def __init__(self) -> None:15super().__init__(16instructions="Ask for the account number, then read it back to confirm."17)1819async def on_enter(self) -> None:20# Callers pause mid-number. Wait longer before ending their turn.21# Takes effect on the caller's next turn.22self.session.update_options(endpointing_opts={"min_delay": 1.5, "max_delay": 5.0})23await self.session.generate_reply(instructions="Ask for the account number.")2425async def on_exit(self) -> None:26self.session.update_options(endpointing_opts={"min_delay": 0.5, "max_delay": 3.0})
Scenario 2: A caller reads out an order number and the transcript comes back garbled
The caller says "K7 dash 3R9Q." The transcript says "k7 3 are 9 queue." What happens next is where the platforms split.
Retell hands the transcript to the model. You can boost keywords in the STT, which helps, but the model still has to guess the order number, call your lookup tool with its guess, and ask the caller to repeat when the lookup fails. Fixing the transcript before the model sees it means taking over generation with a custom LLM server.
Vapi works the same way. Keyword boosting on Deepgram improves recognition, but the model still receives whatever the STT produced and does the guessing. Transforming the turn before the model means hosting a custom transcriber or a custom LLM server that Vapi streams through.
LiveKit edits the turn inside the agent process. You already know the caller's phone number, so you know their open orders. Fuzzy-match the garbled transcript against that list, then replace the turn with the result: "Order K7-3R9Q, shipped Tuesday, arriving Thursday." The model never guesses, never repeats a digit string, and the caller never hears "can you say that again?" No network hop, no second server.
1import difflib2import re34from livekit.agents import Agent, ChatContext, ChatMessage56def normalize(text: str) -> str:7return re.sub(r"[^A-Z0-9]", "", text.upper())89class OrderStatusAgent(Agent):10async def on_user_turn_completed(11self, turn_ctx: ChatContext, new_message: ChatMessage12) -> None:13# Runs after STT and before the LLM, in the agent process. No network hop.14orders = self.session.userdata.open_orders # looked up by caller ID at call start15spoken = normalize(new_message.text_content or "")16candidates = {normalize(o.id): o for o in orders}17match = difflib.get_close_matches(spoken, list(candidates), n=1, cutoff=0.6)18if not spoken or not match:19return # nothing order-like; the model handles the turn as-is2021order = candidates[match[0]]22# Replace the turn. The model relays a verified status instead of guessing digits.23new_message.content = [24f"What's the status of order {order.id}? "25f"[Verified from the caller's account: {order.status}, arriving {order.eta}]"26]
Scenario 3: Handle whatever picks up an outbound call: a person, voicemail, a phone menu, or nothing
Leaving a message is one field on every platform. Telling a menu from a mailbox, getting through the menu, and logging why 30% of last night's dials didn't connect is where they split.
Retell runs two detectors. Voicemail can hang up or leave a message; the IVR detector's only action is hang up, and navigating the menu is a separate feature you turn on. The two are handled separately, so a call classified as IVR never triggers the voicemail action
1// PATCH https://api.retellai.com/update-agent/{agent_id}2{3"voicemail_option": {4"action": {5"type": "prompt",6"text": "Leave a short message for {{name}} asking them to call us back at {{callback_number}}."7}8},9"ivr_option": { "action": { "type": "hangup" } }10}1112// To navigate a menu instead of hanging up, leave ivr_option unset and give the13// Retell LLM a press_digit tool. PATCH https://api.retellai.com/update-retell-llm/{llm_id}14{15"general_tools": [{ "type": "press_digit", "name": "press_digit", "delay_ms": 1000 }]16}1718// The call webhook reports one reason per dial:19// "disconnection_reason": "voicemail_reached" | "ivr_reached" | ...20// A call classified as IVR never triggers the voicemail action.
Vapi picks a detection provider and returns one verdict, voicemail or not. However, a menu, a full mailbox and silence all read as "not voicemail," so the agent starts its pitch. Getting through a menu depends on the model deciding to call the dtmf tool.
1// PATCH https://api.vapi.ai/assistant/{assistant_id}2{3"voicemailDetection": {4"provider": "vapi",5"backoffPlan": { "startAtSeconds": 2, "frequencySeconds": 2.5, "maxRetries": 5 },6"beepMaxAwaitSeconds": 207},8"voicemailMessage": "Hi, this is Acme calling for {{name}}. Please call us back at 555-0100.",9"model": {10"provider": "openai",11"model": "gpt-4.1",12"messages": [13{14"role": "system",15"content": "... If you hear a phone menu, stay silent until the options finish, then use the dtmf tool (for example keys \"w2\") to reach billing."16}17],18"tools": [{ "type": "dtmf" }]19}20}2122// Then POST https://api.vapi.ai/call with phoneNumberId, customer.number and assistantId.23// The end-of-call report gives one verdict: "endedReason": "voicemail", or not.24// A menu, a full mailbox and silence all read as "not voicemail", so the assistant25// starts talking, and getting through the menu depends on the model calling dtmf.
LiveKit classifies with an STT and LLM you pick and a prompt you can edit, returns one of five verdicts (person, voicemail, IVR menu, unavailable line, or uncertain) with the transcript behind it, and in Python starts IVR navigation on its own. Each verdict is a branch in your code and ops can get five reasons a dial failed, not one.
1import os23from livekit import api4from livekit.agents import AMD, AgentSession, JobContext56SIP_TRUNK_ID = os.environ["LIVEKIT_SIP_OUTBOUND_TRUNK"]78async def dial(ctx: JobContext, session: AgentSession, phone_number: str) -> None:9identity = "callee"10session.room_io.set_participant(identity)1112# Start AMD before the callee joins so no early audio is lost. The agent's13# speech is paused until the verdict is in.14async with AMD(session, participant_identity=identity) as detector:15await ctx.api.sip.create_sip_participant(16api.CreateSIPParticipantRequest(17room_name=ctx.room.name,18sip_trunk_id=SIP_TRUNK_ID,19sip_call_to=phone_number,20participant_identity=identity,21wait_until_answered=True,22)23)24await ctx.wait_for_participant(identity=identity)2526result = await detector.execute() # STT and LLM you pick, prompt you can edit27log_dial_outcome(phone_number, result.category, result.transcript) # your ops report2829if result.category in ("human", "uncertain"):30pass # the paused reply plays and the conversation continues31elif result.category == "machine-ivr":32pass # IVR navigation already started (ivr_detection=True is the default)33elif result.category == "machine-vm":34handle = session.generate_reply(35instructions="Leave a brief message asking the customer to call back."36)37await handle.wait_for_playout()38ctx.shutdown("voicemail")39elif result.category == "machine-unavailable":40ctx.shutdown("mailbox full or not set up") # retry later
Scenario 4: Let a human screen the transfer
All three can brief the human before connecting. The difference is what happens after the human hears the briefing.
Retell has two modes. Warm transfer speaks a one-shot whisper, then bridges; the human can't reply. Agentic warm transfer hands off to a separate screening agent you build, which talks with the human and then bridges or cancels. A cancel or timeout sends the caller back to your agent, but your agent never learns what the human said.
1// On the Retell LLM. PATCH https://api.retellai.com/update-retell-llm/{llm_id}2{3"general_tools": [4{5"type": "transfer_call",6"name": "transfer_to_specialist",7"description": "Transfer the caller to a billing specialist after they confirm.",8"transfer_destination": { "type": "predefined", "number": "+14155550100" },9"transfer_option": {10"type": "warm_transfer",11"agent_detection_timeout_ms": 30000,12"on_hold_music": "relaxing_sound",13"show_transferee_as_caller": true,14"private_handoff_option": {15"type": "prompt",16"prompt": "Whisper a one-sentence summary of the caller's issue to the specialist."17},18"public_handoff_option": {19"type": "static_message",20"message": "You're now connected."21}22}23}24]25}26// Warm transfer: the whisper is one-way. Agentic warm transfer adds a screening agent.27// If nobody picks up within agent_detection_timeout_ms, the tool fails and28// your agent carries on with the caller.
Vapi has eight transfer modes. Seven are one-way: two blind, four that speak a fixed message or a generated summary, and one that runs your TwiML on the specialist's leg, then bridge. For a human who can ask a question or decline, you switch to the eighth, warm-transfer-experimental, and prompt a second assistant to run the handoff; if the human declines or doesn't answer, the caller comes back to your assistant or the call ends, per your fallback plan.
1// In the assistant's model.tools, or POST https://api.vapi.ai/tool2{3"type": "transferCall",4"destinations": [5{6"type": "number",7"number": "+14155550100",8"description": "Transfer to a billing specialist after the caller agrees",9"transferPlan": {10"mode": "warm-transfer-experimental",11"transferAssistant": {12"firstMessage": "Hi, I have a caller with a billing question on hold. Can you take it?",13"firstMessageMode": "assistant-speaks-first",14"maxDurationSeconds": 120,15"silenceTimeoutSeconds": 30,16"model": {17"provider": "openai",18"model": "gpt-4o",19"messages": [20{21"role": "system",22"content": "Brief the specialist and answer their questions. Call transferSuccessful when they accept. Call transferCancel on voicemail, no answer, or a decline."23}24]25}26},27"holdAudioUrl": "https://example.com/hold.mp3",28"fallbackPlan": {29"message": "I couldn't reach a specialist. I can keep helping you.",30"endCallEnabled": false31}32}33}34]35}36// The other seven modes are one-way. This one runs a second assistant that37// talks to the specialist; on decline or no answer the caller gets the38// fallbackPlan message and either returns to your assistant or the call ends.
LiveKit runs the handoff as one awaited task. The human lands in a private room with the transcript and three tools: connect to the caller, decline with a reason, or flag voicemail. Whatever they choose comes back to your code.
1import os23from livekit.agents import Agent4from livekit.agents.beta.workflows import WarmTransferTask, WorkflowInstructions5from livekit.agents.llm import ToolError, function_tool67SPECIALIST_NUMBER = os.environ["SPECIALIST_PHONE_NUMBER"]89class SupportAgent(Agent):10@function_tool11async def transfer_to_specialist(self) -> None:12"""Transfer the caller to a billing specialist. Only call after the caller confirms."""13await self.session.say(14"Please hold while I check if a specialist is available.",15allow_interruptions=False,16)17try:18# One awaited task: dials the specialist into a private room, plays hold music19# to the caller, briefs the specialist from chat_ctx, and gives them three tools:20# connect_to_caller, decline_transfer(reason), voicemail_detected.21await WarmTransferTask(22sip_call_to=SPECIALIST_NUMBER,23sip_trunk_id=os.environ["LIVEKIT_SIP_OUTBOUND_TRUNK"],24chat_ctx=self.chat_ctx,25ringing_timeout=25.0,26instructions=WorkflowInstructions(27extra="Say who is calling and what they need, then answer questions."28),29)30except ToolError as e:31# No answer, "human agent declined to connect: <reason>", or "voicemail detected".32# The caller is already off hold. The reason goes back to the model as the tool result.33raise ToolError(f"Could not transfer: {e}. Apologize and offer to take a message.") from e3435await self.session.say(36"You're connected with a specialist now. Goodbye.", allow_interruptions=False37)38self.session.shutdown()
Scenario 5: Dial a list of patients tomorrow morning and report who picked up
This is the one where the hosted platforms often offer a better out-of-the-box experience.
Retell and Vapi ship it as a product. Retell's Batch Call takes a list with per-recipient variables, a start time, a call window and reserved concurrency, and shows Sent, Picked Up and Successful in the dashboard. Vapi's Campaigns take up to 10,000 contacts, a schedule window, a concurrency cap, per-contact webhooks and a pre-dial eligibility hook. Neither needs code.
As an open source framework, LiveKit Agents has no campaign product. Each dial is one dispatch against a SIP trunk you bring, and the list, the schedule, the pacing and the report are yours. It’s a real afternoon of work that the hosted platforms save you. What you get for it is no ceiling: pacing, retries and reporting are whatever you write, and concurrency is a plan tier rather than a per-line fee.
3. Will the agent need to work beyond phone calls, on web, mobile, video, devices or WhatsApp?#
All three answer a phone. Retell adds SMS, MMS and a web chat widget as products. Vapi adds SMS through your Twilio number and a chat API. LiveKit is the only one where the model can see video (like a screen-share), a call can have more than two participants, the agent can run on hardware that isn't a phone or a browser, or answer a WhatsApp call. If the roadmap ends at phone plus text, any of the three works. If it includes web, mobile, video or devices, this question decides it.
| Channel | Retell | Vapi | LiveKit |
|---|---|---|---|
| Phone numbers | Buy US or Canada numbers, inbound and outbound, calls to 15 countries. SIP import for the rest | One free US number, inbound only. Outbound means importing a Twilio, Vonage, Telnyx or DIDWW number, or a SIP trunk | One free US local number per plan, inbound only today. Outbound and other countries run on a SIP trunk you configure |
| SMS and MMS | A product on Retell or Twilio numbers. Two-way, MMS in, agent can text first or mid-call | Glue over your Twilio account. Customer-initiated only, US to US, 24-hour sessions, no MMS. A built-in tool texts mid-call | Sending via Twilio or others is a function tool. Receiving is a small webhook you host that feeds a text-only session. Any carrier, agent can text first, MMS in as an image the model sees |
| Web and in-app chat (typed, not SMS) | Chat-and-voice widget, callback widget, chat API | Widget (voice or text mode, one at a time) and chat API | Voice-first widget with text pane. Text-only widget: build it from the starter |
| Web and mobile | Browser SDK only, current | Web SDK current. Flutter and React Native last shipped mid-2025; iOS has no tagged release | Ten client SDKs, all pushed regularly |
| Video in (the model sees) | Images, audio or video texted in mid-call (MMS), which the agent can describe | No. Screen share and recording reach the call, not the model | Frames into any vision LLM; live camera or screen share with Gemini Live and OpenAI Realtime (Python) |
| Avatars | None | Tavus (though no longer in the docs) | 16 plugins |
| Devices and robotics | None | Python client for desktop | ESP32, embedded Linux, Jetson; ROS and LeRobot through LiveKit Portal |
| Via automation platforms | Via automation platforms | Voice calls through the WhatsApp Connector (Cloud) | |
| More than two participants | Only through transfer | Only through transfer | Rooms hold any number |
The table is the quick version. Here’s each channel in more detail, and what it takes to support it on each platform.
Which platforms include phone numbers?#
All three, with different limits.
Retell is the only one that sells outbound-ready numbers with no telephony account. LiveKit and Vapi each include a free US number, and both are inbound-only.
Outbound on either means bringing a carrier. Vapi imports the number and handles the rest. LiveKit exposes the SIP layer: you configure trunks and dispatch rules, pick carrier and transport, and dial with CreateSIPParticipant. LiveKit is also the only one that answers WhatsApp voice calls natively, through a Cloud-only connector.
Can the agent send and receive SMS?#
All three, in different ways: Retell as a product, Vapi through Twilio, and LiveKit with a small amount of code.
Retell is the only one with SMS as a product: two-way texting on its own numbers, MMS attachments the agent can read, agent-initiated outbound, and A2P registration handled in the dashboard.
Vapi's SMS is a thin layer over a Twilio account you bring. It sets the webhook, keeps a 24-hour session per customer, and replies. It's Twilio-only, US-only, customer-initiated only, and has no MMS.
LiveKit doesn't ship SMS as a separate product, but adding it takes little code. Sending a text is a function tool that calls your carrier's API, and receiving one is a short webhook that passes the message into a text-only agent session. Because it's your code, it works with any carrier, the agent can text first or mid-call, and an MMS photo goes to the model as an image.
Can customers chat with the agent by typing on a website?#
All three, and each widget now handles voice too.
Retell's widget does text chat and browser voice calls, and Vapi's runs in either voice or text mode, one at a time. Both add a chat API. LiveKit's hosted widget is voice-first with a text pane. A text-only widget is two flags on the agent and a fork of the open-source embed starter. It’s buildable, but not prebuilt for you.
Which platforms have mobile SDKs, and are they maintained?#
All three run in a browser; only LiveKit's mobile SDKs are current.
Retell's client SDK is browser-only, though actively maintained. Vapi lists Web, iOS, Flutter, React Native and Python. The web package is current, but (as of October 2026) Flutter and React Native were last published in mid-2025, the iOS SDK has no tagged release, and the Python client hasn't shipped since 2024.
LiveKit ships ten client SDKs (web, Swift, Android, Flutter, React Native, Unity, Unity WebGL, C++, Rust, ESP32), each usually updated weekly.
Can the agent see video or images?#
Only on LiveKit.
In an STT-LLM-TTS pipeline you sample frames from the user's camera or screen into any vision-capable LLM. With Gemini Live or OpenAI Realtime, live video streams to the model directly (Python). Vapi's web SDK can start a screen share and record the call, but nothing actually reaches the model. Retell has no video plane. Its one media path is MMS during a phone call: images, audio or video the caller texts in, which the agent can describe.
Which platforms support video avatars?#
LiveKit has 16 avatar plugins.
Vapi's Tavus integration has dropped out of its documentation and survives only as credential and voice-provider types in the API. Retell has none.
Can the agent run on hardware, including a robot?#
Only LiveKit.
It runs on ESP32, embedded Linux and Nvidia Jetson, so the SDKs behind a phone agent also power a speaker, a kiosk or a wearable. For robots, LiveKit streams camera, sensor and control data in the same room as the voice agent, and LiveKit Portal syncs frames with robot state, arbitrates control between operator and policy, and bridges ROS and LeRobot.
Can more than two people be on the call?#
Only on LiveKit.
A room holds any number of humans and agents, which is how warm transfer, supervisor barge-in, and multi-agent calls work. On Retell and Vapi a call is one caller and one assistant; a second human only arrives through a transfer, and there's no documented way to add a third party and keep the agent on the line.
4. Is cost at scale a relevant concern?#
Around 1,000 minutes a month, cost alone shouldn’t decide this; all three land in a similar range on platform costs. Above it, they split, and Retell and Vapi run 3.5–4x LiveKit's cost (platform, telephony and observability fees; excludes model costs and carrier charges, which are similar on all three). LiveKit also gives you two levers the others don't: cut the LLM bill to a fraction of a cent with fast open models like Gemma 4 through LiveKit Inference, or self-host the whole stack and pay no platform fee at all.
At the end of the day, most teams want a measurable ROI from their investment in voice. In many cases, the biggest return comes from two things: whether your agent delivers a successful experience to your end users, and whether the platform makes it easier for your developers to build, iterate, deploy, and keep optimizing it, which saves engineering time. At scale, though, cost starts to weigh more heavily, and it’s where the platforms diverge most.
Every voice agent bill has the same four parts: a platform fee for running the agent, model costs for STT, LLM and TTS (or a speech-to-speech model), telephony, and whatever you pay for logs and recordings. The model and carrier costs are roughly the same wherever you run them. The platform fee is where the vendors differ, and it's the part that scales with your minutes.
What you pay the platform#
Below are public list prices as of September 2026. Each assumes you bring your own SIP provider (Twilio, Telnyx, etc.). Totals only sum public plan fees, per-minute session fees, platform SIP fees and observability fees. Excludes model costs (STT, LLM, TTS; Retell bundles STT in its session fee), carrier charges, additional concurrency and enterprise discounts, which are similar or negotiable on all three.
| Monthly minutes | Retell | Vapi | LiveKit Cloud |
|---|---|---|---|
| 1,000 | $55 | $50 | $50 (Ship) |
| 10,000 | $550 | $500 | $145 (Ship) |
| 100,000 | $5,500 | $5,000 | $1,400 (Scale) |
Why the gap opens up#
Retell and Vapi charge a flat per-minute platform fee. It's the same rate at your first thousand minutes and your millionth, so the bill grows in a straight line with usage.
LiveKit Cloud charges a lower per-minute rate plus a plan fee, and each plan includes a block of minutes before per-minute charges start. At low volume the plan fee makes LiveKit cost about the same as the others. As usage grows, the included minutes and the lower rate pull the per-minute cost down, which is why the gap widens with every order of magnitude. Past the point where any vendor's list price applies, all three negotiate, and the more relevant LiveKit option becomes self-hosting, where the platform fee is zero.
Model costs: same providers, different billing#
Bring the same STT, LLM and TTS to each platform and the per-minute cost is similar at launch. It diverges as the agent matures, because prompts grow: more instructions, more tools, a knowledge base.
Retell prices each model per minute and scales the bill up once your prompt passes a token threshold, counting instructions, tool definitions, the transcript so far and knowledge base content. Vapi passes through the provider's per-token price. LiveKit Inference bills per token, charges repeated prompt context at the provider's cached rate, and serves fast open-weight models like Gemma 4 for a fraction of a cent per minute. Switching models is one line of code.
Concurrency#
All three include a concurrency allowance and charge for more. Retell and Vapi sell additional capacity per line per month, and with burst turned on, Retell adds a $0.10/min surcharge to calls over your limit. LiveKit's allowance scales with plan tier and can be raised on Scale, with no per-line fee.
Compliance#
All three support HIPAA workloads, but at different price points: Retell includes a BAA on pay-as-you-go, LiveKit includes HIPAA and region pinning from the Scale plan up, and Vapi sells HIPAA mode as a monthly add-on. Vapi and LiveKit offer EU regions; Retell doesn't publish one.
Self-hosting#
Only LiveKit is open source. Run it yourself and there's no platform fee and no concurrency quota from the vendor; you pay for your own infrastructure and any Cloud-only pieces you keep, like Inference, phone numbers and observability. Retell and Vapi have no self-hosted option.
Bottom line: If volume is low and will stay low, the first three questions decide it. If you're heading past 1,000 or 10,000 minutes a month, expect your prompts to grow, or need the audio path in your own infrastructure, LiveKit's pricing is built for that, and it's the only one of the three you can run without the vendor.
Putting it together: which platform fits which team?#
There's no wrong answer here, only a wrong fit. Map your team against the four questions and one platform usually falls out. For most teams, the deciding one is the second: how often you'll need to go past the defaults.
| If your team… | Start with |
|---|---|
| Is building a product where voice is the experience: language learning, coaching, tutoring, companionship, gaming, a device | LiveKit |
| Has no engineers on the agent after launch, and ops or marketing will own and iterate it | Retell |
| Has a developer who can configure and wire up webhooks, the workflow fits a standard pattern, and volume is moderate to low | Vapi |
| Has engineers who will keep iterating, and needs to tune, extend or scale beyond the defaults | LiveKit |
| Needs web, mobile, video, devices or more than two people on a call | LiveKit |
| Expects heavy usage, or needs the audio path in its own infrastructure | LiveKit |
Choose Retell if the agent needs to run without engineers. It's the only one of the three where the whole agent lives in a visual editor a non-developer can own, and it ships the most as products: numbers, SMS, batch calling, a knowledge base. The ceiling is whatever the editor exposes.
Choose Vapi if a developer prefers to assemble the agent from configuration, the defaults are good enough, and the use case follows a well-worn path: simple scheduling, qualification, campaigns. Vapi keeps adding rails for those patterns. You're opting into them rather than writing and optimizing the behavior for your own use case.
Choose LiveKit if your developers would rather build the agent in code than in configuration. With LiveKit, the agent is a program in your repo, so changes go through pull requests, code review, tests and CI like the rest of your stack, and coding agents like Claude Code, Cursor and Codex can build and change it directly. That matters most once you go past the defaults (which most production agents do). It's also the stronger fit if cost at scale matters. The tradeoff is that there's no visual editor or pre-packaged outbound calling campaign product.
One more cut that's easy to miss: Retell and Vapi are built for customer-facing calls. Their feature sets (transfers, voicemail detection, campaigns, SMS follow-up, batch dialing) are primarily contact-center features, and the unit of work is focused around a phone call with one caller and one assistant. If you're building a language-learning app, a voice coaching product, a tutoring or companionship experience, an in-game character, or a voice interface for a device, the product is the conversation itself and LiveKit is the only one of the three built to support that.
Frequently asked questions#
Which platform is fastest to set up: Retell, Vapi, or LiveKit?
All three can have a phone agent answering calls within a day. Retell is fastest for a team with no developer, since the whole agent is built in a visual editor. Vapi and LiveKit both take a developer an afternoon: Vapi via a config and a webhook, LiveKit via a ~30-line agent deployed to LiveKit Cloud. For an engineering team the difference is hours, not weeks; what separates the platforms is how quickly you can change the agent after launch.
Which platform gives you telephony out of the box?
All three. LiveKit is often assumed to require a carrier, but LiveKit Cloud provisions US local and toll-free numbers directly, includes a free local number on every plan, and answers inbound calls without a Twilio account. Outbound calls, transfers and international calls run over a SIP trunk from a provider you choose, and LiveKit is the only one of the three that answers WhatsApp voice calls natively. Retell sells outbound-ready numbers with no carrier account. Vapi's free numbers are inbound-only; outbound means importing a number from Twilio, Telnyx, Vonage or a SIP trunk.
How do you deploy, manage, and host LiveKit?
Run lk agent deploy from your project and LiveKit Cloud builds a container from your code, deploys it, and scales it with call volume. New versions roll out without dropping calls: new sessions go to the new build while active ones finish on the old one. Deployment logs, session metrics, recordings and traces are in the dashboard, with instant rollback and separate non-production deployments on paid plans. If you'd rather run the agent yourself, the same code deploys to Kubernetes, Render or any container platform, connected to LiveKit Cloud transport or a self-hosted LiveKit server.
Which platform has the lowest latency?
None of the three publishes a comparable end-to-end benchmark, and most of a voice agent's latency comes from the models: how fast the STT finalizes, how long the LLM takes to first sentence, how quickly the TTS starts speaking. All three let you pick fast models. Where they differ is what you can do about the rest. Retell exposes two sliders. Vapi exposes the turn model and a set of timing thresholds in config. LiveKit exposes every stage in code: how turns are detected, which turn model runs, endpointing and interruption thresholds, and where the agent runs.
Inference is the other lever. On LiveKit Cloud the agent is hosted next to the media server and LiveKit Inference, so model calls don't leave the cluster. And LiveKit hosts latency-optimized open-weight models, like Gemma 4 31B, on its own GPUs; on LiveKit's public live benchmark it returns a first sentence several times faster than the frontier APIs, and several times faster than the same model served through a general-purpose provider. The benchmark re-runs every ten minutes and shows p95 and p99 next to p50, because a fast median doesn't help the caller who hits the slow tail. If latency is a competitive feature of your product rather than a number to keep under a threshold, it's a difference that matters.
Can you use speech-to-speech models like OpenAI Realtime or Gemini Live?
Yes on all three, as a model choice. The difference is what you can do around it. On LiveKit, a realtime model is one option in the same framework as a pipeline agent, so you can pair it with a separate TTS voice, keep your own turn detection and noise cancellation in front of it, hand off between realtime and pipeline agents mid-call, and (in Python) stream live camera or screen-share video into it. On Retell and Vapi you select the realtime model in configuration and take the platform's handling of turns and audio.
Which platform should you use if voice is the product, not a support channel?
LiveKit. Retell and Vapi are designed around customer-service calls or outbound campaigns: a phone number, a caller, a resolution or a transfer. A language-learning app, a coaching product, a tutoring experience or a voice-driven game has different needs: it lives inside your web or mobile app, sessions run long, the model may need to see video, and there may be more than two participants. LiveKit's client SDKs, video support, multi-participant rooms and device targets are built for that, and the agent is code in your app, so the experience is yours to design.
Which is the most cost-effective at scale?
At a few thousand minutes a month, the three cost about the same on platform fees. Past roughly 10,000 minutes, Retell and Vapi's flat per-minute rates hold while LiveKit's included minutes and lower session rate pull its cost down, so LiveKit runs at a fraction of the others' platform cost by 100,000 minutes. Model costs are similar everywhere, with one exception: Retell scales its per-minute LLM price up as prompts grow, while LiveKit and Vapi bill per token. See question four for the numbers.
Can you self-host any of these?
Only LiveKit. The server, SIP service, egress and agents framework are open source under Apache 2.0, so you can run the whole stack on your own infrastructure and pay LiveKit nothing. What you give up is the Cloud-only layer: managed inference, observability, the dashboard, Agent Builder, Krisp noise cancellation and LiveKit's hosted turn-detection model. Most teams that self-host keep transport on LiveKit Cloud and run only the agent themselves, which removes the per-session fee without losing those features. Retell and Vapi have no self-hosted option.
Can a coding agent like Codex, Claude Code, or Cursor build on these?
All three publish an MCP server and machine-readable docs, so the answer is yes for each. The difference is what the coding agent is being asked to do. On Retell and Vapi, the MCP server lets an assistant drive the platform: create agents, adjust config, place calls. On LiveKit the coding agent writes the agent itself, and LiveKit publishes a docs MCP server, an llms.txt index, a CLI and agent skills that cover workflow design, handoffs and testing, so tools like Claude Code, Cursor and Codex can efficiently build, test, iterate, and deploy an agent end to end. If your plan is to have a coding agent own most of the implementation, the code-first platform is where it has the most to work with.