Introducing Sentiment: How Tango Hears What Customers Feel

Introducing Sentiment: How Tango Hears What Customers Feel

Zach Kamran

Co-Founder & CTO

Garin Kessler

Head of AI Research

Share

TLDR

Simple AI’s Tango model can now tell how a caller is feeling. It scores frustration, confusion, and happiness from the call audio in real time, catching tone a transcript misses. Here’s how sentiment works, why it’s so hard for AI to read, and what it means for conversion, satisfaction, and lifetime value.

Every customer conversation carries two kinds of information: what the customer says, and how they feel while they’re saying it. People pick up on the second kind without trying, but AI does not. We can hear the hesitation in a “sure,” the strain in a “that’s fine,” and the relief when a problem finally gets solved, and adjust ourselves accordingly. AI agents typically work off of transcripts, which omit sentiment entirely. That’s why an agent can answer every question correctly and still feel tone-deaf: it never noticed that the customer was confused, frustrated, or ready to move forward.

Our customers have been asking for a way to know how their callers feel, and we wanted our agents to know, too. So we built it into our most recent model launch.

Today, we’re introducing sentiment in Tango, Simple AI’s conversational awareness model. Tango now tracks three signals throughout every call (frustration, confusion, and happiness), reading them from the audio, as the call happens, with no transcript in the loop.

Why a Transcript Can’t Capture Sentiment

When we introduced Tango, we explained how most voice AI works: transcribing what the customer says, generating a response from that text, then converting it back to speech. Tango moved the decisions that matter most, like end-of-turn and interruption, out of that pipeline by reading them straight from the audio. Sentiment is the clearest example of why that matters.

Much of how we read emotion lives in the parts of speech a transcript throws away: pitch, pacing, volume, emphasis, and how long someone pauses before they answer. Text-based sentiment tools have to infer feeling from word choice alone. That works when the words themselves are emotional, like “this is ridiculous.” It breaks down when the words are neutral and the tone carries the meaning, which is common in customer conversations, where people tend to stay polite even when they aren’t happy. “Thank you” can be gracious, earnest, or deflated, and a transcript records all three the same way. Tango hears the difference.

Why Sentiment Is a Hard Problem

For people, reading emotion is second nature. For machines, it has been one of the harder problems in speech research. There are a few reasons for this.

The first is that there’s no agreed-upon way to divide up emotion. Some research datasets label frustration and confusion. Some measure intensity: how worked up someone sounds, regardless of the feeling behind it. Some split happiness into joy, delight, and satisfaction. Each dataset carves up the emotional space differently, so results from one rarely translate to another.

The second is data. Labeling emotion from audio is slow and subjective, so many datasets rely on actors reading scripted lines, which keeps them small and doesn’t always sound like a real phone call. Even trained labelers don’t always agree on what they’re hearing.

The third is that the field is still young. Recent progress has come from feeding audio directly into language models and classifiers, the same family of approaches Tango is built on.

How Tango Scores Sentiment

Three signals, each scored from 0 to 3

Rather than try to label every possible emotion, we focused Tango on three signals that shape a customer conversation, and ones that an agent can actually act on. Confusion tells an agent to slow down and clarify. Frustration tells it to change its approach. Happiness tells it the conversation is going well.

Each signal is scored on a scale from 0 to 3. 0 means the feeling isn’t present, while 3 is the strongest expression of it: for frustration, think of a caller who’s raising their voice. The levels in between capture how strongly the feeling comes through. Each signal has its own dedicated output in the model, so the three scores move separately, and a caller can register as both confused and frustrated at the same time.

Scores that follow the conversation

Tango updates its sentiment scores continuously as the caller speaks, updating every 80ms. The scores describe the caller’s state over time, rather than tracking isolated outbursts.

Take a caller who was unhappy with their order. As they explain the problem, frustration rises to a 3, and stays there for as they continue talking. When the agent finds the order and offers a resolution, you can hear the relief in the caller’s voice, and the score comes back down.

Tango reacts to the conversation as it happens, and doesn’t try to decide whether a later rise is the same confusion or a new one. It reports how the caller sounds at every point in the call, which is exactly what an agent needs to decide what to do next, and what a reviewer needs to see where a call turned.

One model, one pass over the audio

Sentiment isn’t a separate system added on top of Tango. The same model that detects end-of-turn and interruptions produces the sentiment scores, in the same pass over the same audio. There’s no second pipeline to run, no audio to reprocess, and no added latency for the agent. Because Tango already computes sentiment on every call, turning it on is simply a matter of deciding what to do with it.

In development, we found that training these skills together also made Tango better at each of them. Understanding when a caller is frustrated helps the model understand their pacing, and understanding pacing helps it read frustration.

And because Tango is built for real phone audio, sentiment works on the compressed 8kHz telephony that most customer calls still run on.

What This Looks (and Sounds) Like

The best way to understand sentiment is to hear it. In each example below, listen to the audio, then watch how Tango’s scores move.

Happiness is as useful a signal as frustration. Here, Tango hears a caller’s mood lift, and reacts accordingly.

Confusion rises as a customer tries to identify that the parts he received were the right ones for his car. The agent is patient, communicates what he's looking for, and provides the confirmation he was looking for.

A preview of what’s coming: Tango hears frustration building, and the agent slows down to acknowledge it before moving on.

What Sentiment Makes Possible

A sentiment score on its own is information. Its value comes from what agents and teams do with it, and that falls into two broad categories: changing how an agent behaves during a conversation, and understanding customers better across many conversations.

During the call: agents that adapt

A good human agent adjusts constantly based on how the customer sounds. Sentiment gives an AI agent the same input, so its next move can depend on how the customer feels, not just what they said.

  • Clarifying when confusion rises: Instead of moving on to the next step, the agent can slow down, rephrase, or check that the customer is following.

  • Changing course when frustration builds: The agent can acknowledge the frustration, offer a different path, or bring in someone from your team when a person is the right answer.

  • Matching the customer’s tone: A brisk, cheerful caller and a hesitant one need different pacing. Sentiment lets the agent tell them apart.

  • Recognizing the right moment: When a customer is happy with how the conversation is going, that’s often the natural time to suggest what comes next, whether that’s an upgrade, an add-on, or scheduling the next appointment.

After the call: insight across every conversation

With Tango, every call becomes a record of how the customer felt from start to finish. That means anyone reviewing it can see exactly where the conversation turned, for better or worse, without listening to the whole thing.

Across many calls, the patterns get more useful. If frustration rises every time a particular policy, product, or step in the process comes up, that’s not a problem with one call. It’s a signal about the customer experience itself, and it shows the business exactly where to improve.

We use the same signal to improve our own agents, finding the moments where callers get stuck and fixing them together with each customer.

What This Means for Your Business

Sentiment connects directly to the numbers your team is already measured on.

  • Conversion: A customer who’s confused about price, timing, or next steps is less likely to commit. An agent that hears confusion can clear it up while there’s still a sale to make.

  • Customer satisfaction: Surveys reach a small share of customers, and only after the call is over. Sentiment gives you a read on every call, as it happens.

  • Loyalty and lifetime value: Relationships are built over many conversations. Knowing how each one actually felt to the customer, not just whether the issue was resolved, shows you which relationships are healthy and which need attention.

What’s Next for Sentiment

Sentiment is live in Tango today, running on the same production traffic as end-of-turn, with scores available for every call. Next, we’re putting those scores to work: agent behavior that responds to how customers feel, and sentiment analytics across every call.

Think about the best sales associate you’ve ever had. They knew your name, remembered what you bought last time, and could tell from your voice whether today was a good day. That kind of attention used to be reserved for luxury retail, because it was too costly to give every customer. Sentiment brings us one step closer to offering it on every call.

Want to hear it for yourself? Schedule a demo with our team.

Put your best rep on every call.

Frequently asked questions

What is sentiment in Tango?

Sentiment is a feature of Tango, Simple AI’s conversational awareness model, that tracks how a caller feels during a call. It scores frustration, confusion, and happiness directly from the call audio in real time. Agents can use the scores to adapt mid-conversation, and teams can use them to see where calls turned.

How is Tango’s sentiment different from text-based sentiment analysis?

Text-based tools guess at feeling from word choice alone, so they miss tone when the words are neutral. Tango reads sentiment straight from the audio: pitch, pacing, volume, emphasis, and pauses. That lets it tell a gracious “thank you” from a deflated one, even though a transcript would record both the same way.

What emotions does Tango detect?

Tango tracks three signals: frustration, confusion, and happiness. We picked these because an agent can act on each one. Confusion means slow down and clarify. Frustration means change approach. Happiness means the conversation is going well. Each signal is scored from 0 (not present) to 3 (strongest expression).