Voice AI’s biggest challenge is timing. For a conversation to feel natural, an agent has to recognize the moment a customer is done speaking and craft their response quickly and accurately. Most models still struggle with this, either interrupting mid-sentence, or pausing too long to confirm their turn to speak. Customers are left with conversations that feel robotic rather than human. It’s the biggest obstacle standing between voice AI and mainstream adoption.
We tried all of the end-of-turn models on the market and couldn’t find one that actually worked well in production. In Simple AI fashion, we chose to train our own.
Today, we’re excited to announce Tango, the system behind Simple’s most natural voice agents yet. Tango combines low latency, precise turn-taking, and real-time sentiment understanding to create conversations that feel genuinely human.
How Voice AI Used to Work
Most voice AI models on the market use a cascaded speech → text → speech approach: transcribe what the customer says, generate a response from that transcript, then convert the response back to audio. That transcript has to carry everything — whether the customer's finished talking, how they're feeling, what tone to respond in — and a lot gets lost in translation. The result is both slower and less natural: no matter how fast the agent is, it still has to run the full pipeline before it can say anything.
Because Simple Tango analyzes sentiment from audio, it can pick up on tonal cues before the transcript is finished processing. Where a transcription would see the end of a statement, Tango can hear that the customer is still thinking. More context means quicker, more accurate responses.

Benefits of Tango’s Audio Processing
Improved end-of-turn accuracy
Tango isn’t just faster than competitor end-of-turn models, it’s also more accurate. This is crucial, since false positives mean more interruptions and false negatives mean uncomfortable pauses in conversation. With improved end-of-turn, conversations flow more freely.
Hear what text can’t show
Type “thank you” and there’s no way to tell if it was said with a smile or through gritted teeth. Text strips out tone, prosody, and pacing: the exact signals that make an interaction feel human. Audio-native processing keeps that signal intact, so Tango picks up on sentiment a transcript-based model simply can't.
Speed, without the translation loss
The clearest gain with audio-first processing is speed. The old pipeline adds roughly one second of latency before your agent can even start thinking. Every response waits in that same line. Tango skips it: audio comes in, decisions come out, with both agent and customer channels processed at once. Transcripts are still generated after every call for review and reporting — they just don't sit in the critical path anymore.
Built for the phone line you already have
Most customer service and sales calls happen over the phone, not the internet. As a result, most of the audio a voice AI hears is compressed, 8kHz telephony audio rather than the cleaner 24kHz of a VoIP call. Tango handles both. It plugs directly into your existing inbound and outbound calling systems, without compromising the calls that make up most of your real-world volume. No upgrade needed.
Tango’s Single-Model Architecture
Tango handles both sentiment and end-of-turn detection in a single pass over audio stream, which makes Tango simpler and cheaper to train and deploy. It's also better: training sentiment and end-of-turn detection together, rather than separately, improves performance on both. In our development, we found a customer's tone and rhythm aren't separate problems. One model handling both at once, the way people naturally do, produces the most authentic results.
What’s Next for Simple Tango
Tango is already live for Simple's existing customers, delivering lower latency, higher accuracy, and more intuitive conversations. We'll cover each of these in more depth over the coming weeks, and we're planning an upcoming launch of Tango Mini: a public, open-source version of Tango available to anyone who wants it.
Want to see Tango live? Hit the button below.



