
Just a few weeks ago, Simple founder and CEO Catheryn Li spoke at Genesys Xperience about what actually separates voice AI agents that work from ones that just sound good in a demo. For those of you who missed it, here are a few of our key takeaways.
Voice AI comes down to two things: infrastructure and behavior
Cat split the problem of building voice AI into two broad categories. The first was infrastructure, the engineering that happens behind the scenes and that most people never think about directly. This is what determines whether a conversation feels natural. Does the voice sound human? Does it know when it is actually its turn to speak? How quickly does it respond, and can it handle being interrupted without falling apart?
In our system, every single turn of conversation runs multiple models in parallel: speech to text, voice activity detection, end-of-turn detection, interruption detection, and noise cancellation. That output feeds into an inference layer that decides what knowledge to access, what guardrails to apply, and what the agent should say next, after which a text-to-speech model turns the agent’s response back into audio.
If you were to chain together off-the-shelf APIs for this task, you would likely end up with response times of one to three seconds per turn—and a stilted, robotic conversation experience. At Simple, we train and self-host our own models across the stack, and we are proud to have reduced turn latency to an industry-leading 550 milliseconds.
The behavior side is where businesses have real control
The second half of the equation is behavior, or what your agent actually says and does. Cat framed this in three tiers:
Does the agent route and triage correctly?
Does it answer accurately?
Does it take the right action when it matters?
Answering accurately depends on the agent knowing how and when to pull from the right knowledge base, especially when a company has several disconnected ones. Taking the right action is equally crucial, since placing an order incorrectly creates significantly more work than doing nothing at all.
Good documentation matters, but so do your call recordings
The better and more complete a company’s SOPs and knowledge base, the better the results. This is not surprising on its own, but what is less obvious is that even strong documentation only shows a fraction of the picture inside a contact center. We ask clients to let us train on past call recordings and transcripts so we can learn from how their best agents actually handle situations the SOPs do not cover. Those recordings also become the basis for evals, where we build customer personas to continuously test that the agent handles real scenarios correctly as it evolves.
A different way to think about improving customer experience
One of the more interesting points from the talk was a mindset shift. Traditionally, contact center leaders spend a lot of energy coaching their weaker reps to close the gap with their best ones. With AI, the question becomes different: how do you get the system to perform as well as your best rep? Once that is done, how do you make your best rep even better? AI also makes A/B testing in call centers quick and effortless. Instead of taking months to roll out a change across a human team, you can test a new message, tone, or even voice and see results tomorrow.
The results
Across deployments, Simple AI sees an average resolution rate of 70 percent, along with higher conversion and upsells compared with live reps. We use the word resolution intentionally, rather than containment or deflection. At Simple, the goal is not to get customers off the phone. It is to solve their problems.




