Serve AI provides AI-powered phone agents for field service companies, handling inbound customer calls that would otherwise go to a contact center. One of its largest customers is a pest control company doing about $150M a year that was spending $8M on a 100-person contact center. Serve now handles every inbound call, including account lookups, eligibility, scheduling, billing, and cancellations, with calls transferred to a human when needed.
Serve originally built the agent on Retell, where behavior lives in a conversation flow graph. For that one customer, the graph reached 224 nodes and over 600 edges, and a single change had become a day of work. Anthony Rassi, Serve's CTO, decided to collapse the whole thing into one prompt on LiveKit, with the right procedure injected at each turn. That only became viable once smaller, faster models caught up. Anthony had read LiveKit's write-up on latency-optimized voice agents with Gemma and wanted to see whether a single-prompt agent on a model that fast could still hold the behavior the graph had been enforcing.
The hard part wasn't writing the new agent. It was knowing whether it did everything the old one did, and could meet Serve's evolving needs.
"When you go from one agent architecture to another, you want to make sure the second works at least as well as the first, ideally better. The way of knowing that is Agent Simulations. We ran about a hundred real production calls through it and got to a 98% pass rate before we cut anyone over."
Anthony Rassi, CTO, Serve
Anthony had been running his evals on Hamming, and had also looked at Coval and Bluejay. For his setup, each required more integration work than he wanted, including connecting external test harnesses into LiveKit rooms before he could start evaluating the agent. What he needed was specific: did the call end in the right state, such as an appointment booked for the correct Wednesday? Did the agent confirm the date before booking it? And how long did the caller wait for a response? Agent Simulations let him define and run those checks directly against the agent.
Anthony picked about a hundred production calls the Retell agent had handled, used Claude to turn each into a simulation scenario, and ran the rewritten agent against the set. Over ten to fifteen iterations, he adjusted the prompt until the suite reached a 98% pass rate and median turn latency came in at 1.3 seconds, against a 1.5 second ceiling. Then he shipped. The first account took a week.
In production, the rewritten agent did more than reach parity with Retell: it increased containment by 15%, which Serve estimates saves the customer an additional $320,000 a year.
The improvement came from the new architecture. In the old graph, a caller at the booking node who suddenly raised a billing question might have no defined path to follow, causing the agent to improvise. With a single prompt, the full procedure is available at every turn, so the agent can handle those shifts more reliably.
Agent Simulations gave Serve the confidence to make that rewrite in the first place. The team could test the new agent against real production calls, iterate until it reliably matched the expected behavior, and then cut traffic over. The first migration took one week, with no rollback, and the new agent ultimately performed better than the one it replaced.