A customer was asked how a product they just bought is working out. Here is exactly what they said, word for word. Read it first, then listen.
It's fine. It does what it's supposed to.
The recording. Same words you just read.
A different customer, asked the same thing, gave the exact same words. Here is that recording.
Same sentence, word for word. On the page these two answers are identical: any transcript would file them as the same response. Out loud they are opposites. The whole difference lived in the delivery, which is the part a transcript throws away.
A Voice Emotion AI scores the first, flat answer, and an obvious one for comparison.
It's fine. It does what it's supposed to.(the flat one)
Satisfied · 91% sure
The obvious one:
Angry · 95% sure
It is 91% sure the flat answer is satisfied, and 95% sure the loud one is angry. It is confident both times, and right only about the loud one. The obvious feelings are the ones you never needed a machine to spot. The quiet, mixed ones, which is what most research turns on, are exactly where it stays confident and gets it wrong.
You read past the words yourself the moment you heard the delivery. That second channel is real, and it is worth collecting. What the machine adds is false confidence about the very cases that matter. Analyse what people say, and treat any tone score as a claim to check, never a measurement to report.
The three clips are recorded output from a text-to-speech model (Google Gemini), generated for this widget on 17 August 2026; the first two are the identical sentence read two ways. The Voice Emotion AI scores are an illustrative example of a documented failure mode, patterned on a 2025 expert-checked speech-emotion benchmark (roughly 95% on loud states like anger, roughly 63% on close ones like sadness versus distress). They are not a live model call.