Voice and Multimodal Research: Ask Aloud

Article M1-03

Voice and video capture more than the words a person types. Whether a machine can read what that extra carries is a very different question.

In brief

Every spoken answer has two channels. There are the words, and there is the delivery: the pace, the pauses, the catch in the voice, and on camera the face and the hands. For decades research kept only the first channel, because the audio got typed into a transcript and the transcript was all anyone could analyse at scale. That has changed. A machine can now take in the raw audio and video, run the interview aloud, and transcribe it cleanly for thousands of people at once. But the real gain is that the machine can read the second channel, and tell you how someone felt from the sound of their voice. That claim is unfortunately the weakest part of the whole picture, and the evidence that says so is older and firmer than the evidence that says otherwise. So: collect the voice but do not let a model tell you what that voice means.

What the theory says

The theory

A person answering a question does two things at once. They choose words, and they deliver them in a particular way. The words carry the content. The delivery carries something else, and everyone knows this from ordinary life, because the same sentence can land as sincere, sarcastic or defeated depending on how it is said. Researchers have always known the delivery matters. The problem was that they could not use it. A recorded interview was transcribed, and the transcript, the words alone, was what got coded and counted. The second channel, how the thing was said, was thrown away at any real scale, or it was read by a single analyst working through one tape at a time.

After 2022, that stopped being a hard limit. A large language model that handles audio and video can take the recording itself as input, not just a typed version of it. It can hold a spoken conversation, listen to the reply, and produce a clean, speaker-labelled transcript afterwards. A machine can now run the interview itself, which we cover in AI-moderated interviews. What matters here is not who asks the questions but the form the answer arrives in. When the answer is spoken or filmed rather than typed, two separate things become possible, and keeping them apart is the whole point of this article.

The first is collection. Spoken answers are easier to give than typed ones, so people give more. Höhne, Gavras and Claassen (2024) had smartphone-survey respondents answer sensitive open questions by typing or by speaking, and the voice answers ran more than twice as long and covered more ground. The transcription that used to be slow and expensive is now fast and cheap. Harrington (2023), working with the hard case of police-suspect interviews, showed that automatic speech recognition can carry a real share of the transcription load, though it struggles with overlapping speech and strong accents and still needs a human check. But when the speakers don't overlap, the machine can also easily mark who spoke when, which is the unglamorous part that makes a two-person or group recording usable as data rather than a wall of undivided text.

Video adds its own layer. Conrad et al. (2023) ran survey interviews with video and found it changes both the quality of the data and how the interview feels to the person answering, so the picture channel is not free either way, and it has to earn the added cost and the added intrusion. And the interview can now be spoken at a scale no human staff could reach. Tirumala et al. (2025) assessed AI voice interviewers for research and found they clearly beat the old automated phone systems on holding a real exchange. Ding, Taoka and Nakatani (2026) built short voice-chatbot surveys that run in the field, a spoken cousin of the conversational survey.

The second thing is inference, and it is a different claim entirely. Inference is the machine listening to the delivery and telling you what it means: this person was anxious, that one was excited, this answer was not sincere. The collection gain is well founded. The inference claim is where the marketing has run far ahead of the evidence, and separating the two is the single most useful thing a buyer of this technology can do.

Controversies

The first open argument is whether talking, rather than typing, actually gets you better answers or just longer ones. The intuition is that people open up when they speak. That is only half right. Höhne, Gavras and Claassen (2024) did find voice answers longer and broader, but they found no difference in the emotional intensity of what was said, and there is a repeated pattern in survey research that typed answers can carry more disclosure on awkward topics, because typing is private in a way that speaking aloud is not. So voice changes the answer, but which way better points depends on what you are trying to measure. Whether people tell a machine more than they tell a person is its own question, and we take it up in the machine interviewer effect.

The louder argument is about the second channel. Vendors increasingly sell emotion or sentiment read straight from the voice as a built-in feature, a number attached to every answer saying how the person felt. But the one independent academic assessment of AI voice interviewers found the opposite of the sales pitch. Tirumala et al. (2025) judged the emotion-detection side not good enough for serious use, alongside transcription errors that persist and follow-up quality that is uneven. So the same tool that is a genuine step forward for collection is, on its emotion-reading feature, resting on a claim that the nearest thing to an independent test does not support. The deeper question of whether any machine can read emotion from a face or a voice is taken up in emotion AI on trial. Here the point is narrower and it is about the research decision: the emotion score is the part of the product you should trust least.

Limitations

The clearest limit comes from putting a number on it. Schuhmann et al. (2025) built a careful benchmark for reading emotion from speech, checked by experts. Machines did well on loud, obvious states, scoring around ninety-five percent on something like anger. They did far worse on states that sound similar to each other, scoring around sixty-three percent when the task was to tell sadness apart from distress. That gap is the whole problem in one statistic. The obvious emotions are the ones you did not need a machine to spot. The subtle, mixed, half-hidden feelings are the ones research actually wants to understand, and they are exactly where the reading falls down. Consider the customer who says the product is fine in a flat voice. The interesting reading is whether that flatness is quiet satisfaction, mild disappointment, or someone who has stopped caring, and those are precisely the close, low-key states the benchmark shows machines cannot yet tell apart. A tool that is confident about anger and lost about ambivalence is confident in the wrong places for research.

This limit is not new, and it is not the machine's invention. Human listeners have the same problem. Laukka and Elfenbein (2021) gathered the cross-cultural studies and found that people recognise emotion in a voice better when the speaker shares their culture, and the accuracy drops the further apart the two cultures are. A model trained mostly on one kind of speaker carries that bias forward. It reads the people it saw a lot of during training better than the people it saw rarely, so a system tuned on confident native speakers of one language will tend to read a heavy accent, a second language, or an unfamiliar speech rhythm less accurately, and then report that lower accuracy with the same confidence. There is broad agreement that emotion recognition trained on one set of recordings degrades when it meets a different set, though a clean figure for how much is hard to find. As far as this research could establish, no widely accepted number pins down the size of that drop.

There is a sharper warning from history. The idea that a machine can hear the truth in a voice has been sold before, as voice-stress lie detection, and it did not survive scrutiny. Lacerda (2013) reviewed the commercial claims and found them unsupported, with performance barely above chance and the technology rejected by the scientific and legal bodies that examined it. The vocal signal does carry emotion-relevant information. Juslin and Laukka (2003) established that decades ago, in the sense that pitch, loudness and speed do track feeling to a degree well above chance. But well above chance is a long way from reliable enough to score a customer, and the distance between those two is where careful research lives.

Underneath all of this sits a plainer problem. No independent party has validated the emotion output of any commercial voice tool. The strong numbers come from the companies selling the tool, and a targeted search for outside confirmation of their emotion claims returned nothing on point. The collection gains are real and partly checked by academic work. The inference claims are, for now, marketing.

Open questions

Four things are genuinely unsettled. First, whether the newest general-purpose models that handle audio directly are any better at reading fine emotion than the dedicated systems that came before them, or whether they inherit the same ceiling. Second, whether speaking really buys more candour than typing, or only more words. Third, how far the reading degrades when a tool trained on one population meets another, which everyone agrees happens and nobody has cleanly measured. Fourth, whether the picture channel, the face and the gestures on video, adds real signal beyond the audio or mostly adds cost and risk. None of these has a settled answer, and the field is moving fast enough that the answers will change more than once.

So what

The choice used to be simple, because voice and video were expensive to collect and impossible to analyse in bulk, so most research stayed on the page. That has flipped. You can now run spoken interviews and video diaries at scale, and get clean transcripts back. The gain is real, and it is on the collection side: reach, ease, and richer answers in people's own words. The temptation that comes with it is to buy the other promise as well, the one where the machine hands you a feeling attached to every answer. Treat the first as a genuine new capability and the second as a claim to be checked, and most of the decisions below follow.

For research practice

Use voice and video for what they are good at. Running spoken interviews and video diaries lets you reach people who would never fill in a form, and it pulls longer, fuller answers out of them. Take those gains. What the method can honestly claim today is better collection: more reach, easier participation, and clean transcription at scale (Höhne, Gavras and Claassen, 2024; Tirumala et al., 2025). What it cannot honestly claim today is a reliable read of how someone felt from the sound of their voice (Schuhmann et al., 2025; Tirumala et al., 2025). So analyse the words, treat any automatic emotion score as a flag to check rather than a finding, and keep a human in the loop for the delivery. What happens to the transcripts after collection, the coding and the analysis, carries its own traps, covered in the speed and truth of AI research.

For companies

This makes it cheap to listen to customers in their own voices, at a scale that used to be out of reach. Buy it for that. The feature to be wary of is the emotion or sentiment score read from the voice, because it is the part with no independent validation and the part a vendor is most tempted to oversell. A dashboard that colours every customer call by mood looks powerful and may be measuring very little. If you research your own staff, note that using a system to infer their emotions from voice or video is now banned outright in the European Union under the Artificial Intelligence Act (European Union, 2024), so the compliance question arrives before the accuracy one.

For political parties

The appeal is obvious and the caution is sharper here than anywhere. Reading voters' emotions from the sound of their voices, at scale, is exactly the kind of tool a campaign would want and exactly the kind the evidence does not support. Two things line up in the same direction, which is the useful part. The responsible move and the rigorous move are the same move. An emotion read that is barely better than chance on the subtle states is worthless as intelligence even before it is questionable as ethics, because acting on a feeling the machine mostly guessed is acting on noise. Inferring emotion from voice or face is also already restricted by law in several settings under the Artificial Intelligence Act (European Union, 2024). Use voice to hear what people say and how they put it in their own words. Do not buy a machine that claims to score how they felt.

For government and policy

Consultation in citizens' own spoken words is a real gain, and it can reach people that written forms never do. Two duties come with it. The first is disclosure. Under the Artificial Intelligence Act (European Union, 2024), a person exposed to an emotion-recognition system has to be told it is running, and the profession's own rulebook already treats a voice or video recording as personal data owed consent and care (ICC/ESOMAR, 2025), a duty covered in the 2025 Code. The regulatory detail sits in regulation reaches fieldwork. The second duty is inclusivity. Comfort with speaking to a machine, and the way a voice is read, are not spread evenly across accents, ages and first languages, so a careless rollout will hear some citizens more clearly than others and call the result the public view.

How to use this

Start with the test that contains the whole argument. Take one recorded answer and look at it two ways: first as the plain transcript, then with the audio played alongside. Notice how much your own read shifts once you can hear the hesitation or the warmth. That shift is real, and it is why the second channel is worth collecting. Then notice something else. The richer read you just formed was your inference, built from context you happen to share with the speaker, and it is exactly the judgement a machine does worst when the speaker is nothing like the voices it was trained on. From there, three habits keep you honest. Collect voice and video for reach and for the words, and analyse the words. Treat any automatic emotion or sentiment score as a claim to test against reality, not a measurement to report. And say plainly, wherever a machine did the listening, what it was allowed to conclude and what it was not.

Case studies

The European Union, on where the line now sits. The Artificial Intelligence Act bans using AI to infer a person's emotions from biometric signals, including the voice, in the workplace and in education, and it requires anyone running an emotion-recognition system elsewhere to tell the people exposed to it (European Union, 2024). This is the clearest public verdict on emotion-from-voice inference to date. It is worth reading as a signal in itself: a regulator has drawn a hard line around exactly the capability vendors are selling hardest, which does not happen to a technology whose accuracy is settled.

Outset, on the collection gain. The interview platform Outset markets AI-moderated voice, video and text interviews that run in many languages and with many people at once, and reports very large usage across enterprise clients (Outset, 2026). The usage figures are the company's own and should be read as claims rather than audited facts. But the case makes the collection shift concrete, because the thing being scaled is the gathering of spoken answers, which is the part of this technology on the firmest ground.

Hume AI, on the inference over-claim. Hume AI builds an interface designed to read emotion from the voice and respond to it, and presents its outputs as probabilities grounded in affective science (Hume AI, 2024). The honesty of framing outputs as probabilities is worth noting. So is the gap the rest of this article describes: no independent party has validated that the emotion read is accurate on the subtle states that matter, and the benchmark evidence suggests those are exactly where it struggles. It is the clearest example of the second channel being sold as solved when it is not.

References

Conrad, F.G., Schober, M.F., Hupp, A.L., West, B.T., Larsen, K.M., Ong, A.R. and Wang, T.D. (2023) 'Video in survey interviews: effects on data quality and respondent experience', methods, data, analyses, 17(2), pp. 135–170. Available at: https://doi.org/10.12758/mda.2022.13 (Accessed: 17 August 2026).

Ding, F., Taoka, Y. and Nakatani, M. (2026) 'LLM-based voice chatbot surveys as an alternative to post-experience questionnaires: probe-controlled, ultra-short field interviews', Proceedings of the Design Society, 6, pp. 2253–2262. Available at: https://doi.org/10.1017/pds.2026.10583 (Accessed: 17 August 2026).

European Union (2024) Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Available at: https://eur-lex.europa.eu/eli/reg/2024/1689/oj (Accessed: 17 August 2026).

Harrington, L. (2023) 'Incorporating automatic speech recognition methods into the transcription of police-suspect interviews: factors affecting automatic performance', Frontiers in Communication, 8. Available at: https://doi.org/10.3389/fcomm.2023.1165233 (Accessed: 17 August 2026).

Höhne, J.K., Gavras, K. and Claassen, J. (2024) 'Typing or speaking? Comparing text and voice answers to open questions on sensitive topics in smartphone surveys', Social Science Computer Review, 42(4), pp. 1066–1085. Available at: https://doi.org/10.1177/08944393231160961 (Accessed: 17 August 2026).

Hume AI (2024) Empathic Voice Interface (EVI). Available at: https://www.hume.ai/ (Accessed: 17 August 2026).

ICC/ESOMAR (2025) ICC/ESOMAR international code on market, opinion and social research and data analytics. 5th edn. Available at: https://iccwbo.org/news-publications/business-solutions/iccesomar-international-code-market-opinion-social-research-data-analytics/ (Accessed: 17 August 2026).

Juslin, P.N. and Laukka, P. (2003) 'Communication of emotions in vocal expression and music performance: different channels, same code?', Psychological Bulletin, 129(5), pp. 770–814. Available at: https://doi.org/10.1037/0033-2909.129.5.770 (Accessed: 17 August 2026).

Lacerda, F. (2013) 'Voice stress analyses: science and pseudoscience', Proceedings of Meetings on Acoustics, 19(1), 060003. Available at: https://doi.org/10.1121/1.4799435 (Accessed: 17 August 2026).

Laukka, P. and Elfenbein, H.A. (2021) 'Cross-cultural emotion recognition and in-group advantage in vocal expression: a meta-analysis', Emotion Review, 13(1), pp. 3–11. Available at: https://doi.org/10.1177/1754073919897295 (Accessed: 17 August 2026).

Outset (2026) AI-moderated interviews. Available at: https://outset.ai/platform/interviews (Accessed: 17 August 2026).

Schuhmann, C., Kaczmarczyk, R., Rabby, G., Friedrich, F., Kraus, M., Nadi, S., Kersting, K. and Auer, S. (2025) EmoNet-Voice: a fine-grained, expert-verified benchmark for speech emotion detection. arXiv:2506.09827. Available at: https://arxiv.org/abs/2506.09827 (Accessed: 17 August 2026).

Tirumala, S., Jain, N., Leybzon, D.D. and Buskirk, T.D. (2025) Mic drop or data flop? Evaluating the fitness for purpose of AI voice interviewers for data collection within quantitative and qualitative research contexts. arXiv:2509.01814. Available at: https://arxiv.org/abs/2509.01814 (Accessed: 17 August 2026).

Explore the idea

Let’s talk

Invisible forces shape your world — until you hire Latenta®

Contact