AI-Moderated Interviews: The interviewer who never tires
Article M1-01
A machine can now run a probing interview at survey scale. The question is whether that is real depth or just the look of it.
In brief
For seventy years, research forced a choice. You could run a survey that reached thousands but locked everyone into the same fixed answers. Or you could run a depth interview that understood one person at a time but never scaled. A language model that runs the interview itself, reading each answer and asking its own follow-up, removes the choice. The thing that does the work is the follow-up, not the automation. The early evidence is good on breadth and on matching human interviewers for basic quality. It is thinner on the deep, specific understanding that depth interviews exist for. And the boldest claim, that the answers predict what people do months later, rests on a single study. The research is real. The marketing has run ahead of it.
What the theory says
The theory
The depth interview and the survey have always sat at opposite ends of a trade-off. A survey scales because it is standardised. Everyone gets the same questions in the same order. That makes the answers easy to count, but it also limits them to what the researcher already thought to ask. A depth interview works the other way. It follows the person, asks why, and pushes past the first easy answer to the reasons underneath. That is the whole point of qualitative work, and it has always been expensive, because it needs a skilled human who can only run one conversation at a time. So you could have depth, or you could have numbers. You could not have both. Every research plan was built around that choice.
After 2022, that limit stopped being fixed. A large language model can hold the conversation itself. It asks an open question, reads the answer, and decides what to ask next. Geiecke and Jaravel (2024) built an open-source platform that does this and ran it with thousands of respondents. They argue it closes the old gap between qualitative and quantitative work. Wuttke et al. (2024) tested the same idea in a controlled experiment. They randomly assigned people to be interviewed by an AI or by a human on political topics, and the AI produced data as good as the human interviewer, at a scale no human could reach.
The key point is what does the work. It is easy to assume the breakthrough is automation, that the machine is just a cheaper interviewer. It is not. The breakthrough is that the machine adapts. A static open-text box has been in surveys for decades, and it mostly collects short, tired, useless answers. What changed is that the machine now reads your answer and asks a real follow-up, the way a good interviewer leans in and asks you to say more. Chopra and Haaland (2026) find that AI-led interviews produce richer responses than other survey-based qualitative methods. Their work is the clearest case that the follow-up is the active ingredient. Jacobsen et al. (2025) go further in a controlled study of four kinds of probe. They show that good follow-ups raise the quality of what people say, and that different probes suit different stages of a project. One vendor reports that more than seventy percent of the final insight in its studies comes from the AI's follow-up questions, not the opening one (Conveo, 2026). That is a marketing claim, not an independent finding, but it points at the same mechanism the research describes. The value is in the question the machine asks next, not the questionnaire.
It helps to picture what the machine does that a form cannot. Ask a form why someone switched brands and you get one line: too expensive. Ask an adaptive interviewer the same thing and it hears too expensive, then asks compared to what. It learns the person was comparing against last year's price, not a rival on the shelf. That second fact is the one a marketing team can act on, and it only exists because something followed up. Geiecke and Jaravel (2024) show this holds across very different questions. Their platform held up against human experts on everything from the practical reasons behind a decision to people's views of the wider world. Vendors have turned this into a numbers race, advertising interviews that go five to seven layers deep, or up to ten follow-ups to a single question. That depth is real, and it is easy to measure, because you can count the layers. Whether counting layers is the same as understanding the person is the question the next section takes up.
Controversies
The trouble starts with the word depth, because it is doing two jobs. One meaning is about process: how many good follow-ups the interview asks, how many layers deep it goes. Machines are already strong here, and vendors compete on it, advertising probes five to seven layers deep or up to ten follow-ups per question. The other meaning is about understanding: whether the interview captures why this particular person feels this way, in their own words. This is where the machine is weakest. Cuevas et al. (2025) ran the largest peer-reviewed study of the method so far. They built a scale to measure exactly this richness and found that LLM interviewers hit the usual marks for chatbot quality but rarely captured a person's specific motives or personal examples. The tool sells the first kind of depth. The researcher is paying for the second. They are not the same thing, and a study that only counts follow-ups will not see the gap. A machine can ask four smart questions about why you cancelled a subscription and still miss that the real reason was one rude email from support, because you never quite said it and it never knew to look. A skilled human catches the flicker and pulls the thread. That is the part the metrics do not yet measure.
A second claim is louder in the marketing than in the evidence: that people open up more to a machine than to a person. It is plausible, and there is a name for why. Talking to a machine online can carry a felt anonymity, a sense that no one is really watching, that talking to a person does not. Williams and Ingleby (2025) point to this as one reason an AI probe can draw out answers a human interviewer would not. Decades of survey research point the same way. People give more honest answers about sensitive things when no interviewer is present. Rickwood and Coleman-Rose (2023) found that young people reported lower psychological distress to an interviewer than they did on their own, which is the social-desirability effect at work, and it suggests a machine might loosen tongues. But it is not that simple. Höhne, Neuert and Claassen (2025) found that putting a visible, on-screen agent in front of people can bring the social pressure back, because a face, even a synthetic one, makes people perform. So a machine interviewer might get more candour, or less, depending on how human you make it. Nobody has settled the question for conversational AI. Vendors assert the advantage. The research has not confirmed it.
The boldest claim is that these interviews predict behaviour. Chopra and Haaland (2026) report that the answers people give an AI interviewer are internally consistent and go on to predict what those people actually do six months later. If it holds, it matters, because it moves the method from producing interesting transcripts to producing something you can bet a decision on. But it rests on that one study. A dedicated search for other work linking AI-interview answers to later behaviour found nothing. It is a strong, specific result from a single working paper, not yet reproduced, and it should be read that way.
Limitations
The richness gap is the first hard limit, and it is dangerous because it hides. An interview can pass every generic quality check, run long, feel smooth, and still miss the one specific reason that would have changed the decision. This is the same failure that shows up when AI analyses qualitative data instead of collecting it: the output looks reasonable and is wrong underneath. The companion article on speed and truth in AI research covers that case. A number that looks fine is the hardest kind of error to catch.
Rapport is the second limit. In human interviewing, the bond between interviewer and respondent is not decoration. It shapes the data. Horsfall et al. (2026), working with clinical interview data, found that rapport, and the data quality it supports, depends partly on how similar the interviewer and respondent are in age, education and background. A machine has none of that to share. It can be endlessly patient and never tire, which is a real advantage. But whether it can build the kind of trust that gets a guarded person to say the true thing is still an open question.
Then there are the risks that come from the method working too well. The same easy manner that gets people talking also gets them oversharing. Li, Xiao and Li (2026) built privacy controls for interview chatbots exactly because people disclose more to them than they mean to, which turns a data-collection win into a duty-of-care problem. There is also an inclusivity gap. Williams and Ingleby (2025) note that the felt anonymity of an AI probe helps some people speak freely, but comfort with being interviewed by a machine is not spread evenly, so a careless rollout can quietly shut out the people who are hardest to reach anyway. And most commercial systems run on proprietary models and hand every task to the model at once, which makes them hard to reproduce and hard to audit. Gårdhus et al. (2026) built an open-source alternative that keeps the wording and order of questions under human control. That control is not a technicality. If the model rewords the question each time, two respondents were not asked the same thing, and the comparison the study depends on quietly breaks down. There is a data-security cost too: sending thousands of candid, personal interviews through a third party's model puts sensitive words somewhere the researcher does not fully control.
There is a simpler problem under all of this. No commercial platform has been tested by anyone independent. Every number they publish, on depth, on scale, on how much people prefer the machine, comes from the company selling it. The research is good, and it shows a real but limited tool. The marketing shows a finished one. Nobody has closed the gap between the two yet, and that gap is what a buyer should watch.
Open questions
Four things are genuinely unsettled. First, whether a machine can build real rapport, or can only ask the questions a rapport-rich interview would have asked. Second, whether the six-month behaviour prediction holds up outside the study that first reported it. Third, who gets left out when the interviewer is a machine, and whether that skews the findings. Fourth, how deep the understanding can really go, given that today's systems clear the generic bar and miss the specific one. None of these has an answer yet, and the field moves fast enough that the answers will change more than once.
So what
The choice between depth and scale used to be something you built your research around. Now it is a setting you can change, and that has real consequences for what you can commission and what it costs. But the depth a machine gives you is not yet the depth the word promises, and the strongest claims are young and mostly made by the people selling the tool. So treat AI-moderated interviewing as a strong tool for breadth with a known blind spot in the middle, and design around that blind spot rather than pretend it is not there.
For research practice
Use it for what it is good at: scaled exploration. Running hundreds or thousands of open-ended conversations to find out what themes are out there, in people's own words, is something you could not do before, and it is worth doing. But judge the tool on the quality of its follow-ups, not the number of interviews it ran, because the follow-up is where the value is. Keep a set of human-moderated interviews as a holdout, and read them against the machine's transcripts to see what it missed. Here is what the method can honestly claim today. It matches human interviewers on basic quality and is strong on breadth of themes (Wuttke et al., 2024; Geiecke and Jaravel, 2024). It is weaker at capturing specific personal motives (Cuevas et al., 2025). And its ability to predict later behaviour is one promising study, not a settled fact. What happens to the transcripts afterwards, the coding and analysis, is a separate question with its own traps, covered in the companion article on AI and the speed of research.
For companies
This makes continuous discovery cheap and always on. You can listen to customers at a scale and speed that used to be impossible, in their own words rather than through a rating scale. The risk is mistaking coverage for understanding. A thousand shallow conversations that all point the same way can feel like strong evidence and still miss the real reason people behave as they do, which is the exact thing the machine is weakest at capturing. Before you bet a product decision on it, check it against a smaller human-run study, and treat the AI interviews as the wide net, not the final word.
For political parties
The appeal is obvious. You can ask thousands of voters not just what they think but why, in open conversation, and get back something richer than a tracker poll. The research supports the basic capability on political topics (Wuttke et al., 2024; Geiecke and Jaravel, 2024). Two warnings matter more here than usual. First, the mode changes what people will admit. A machine may draw out sensitive views a human canvasser never would, or it may flatten them, and you cannot assume which. Second, the machine cannot read the room the way a good organiser does, so the warmth and judgement that turn a doorstep chat into a persuadable voter are not in the transcript. Use it to understand the electorate, not to replace the human contact that moves it.
For government and policy
Consultation at population scale, in citizens' own words, is a real democratic gain. It could be the difference between a policy shaped by the few who fill in forms and one informed by the many who will actually talk. Two duties come with it. The first is inclusivity. Comfort with being interviewed by a machine is not spread evenly, and a rollout that ignores this will overweight the confident and the digital and under-hear everyone else. The second is a duty of care. The same quality that gets people talking gets them oversharing (Li, Xiao and Li, 2026), so informed consent and real data protection are part of the method, not paperwork added at the end. A public body using this owes people clarity about what they are talking to and what happens to what they say.
How to use this
Start with the test that contains the whole argument. Take one open question and ask it two ways: once as a static text box, once through an adaptive probe that follows up. Then count the distinct themes each one surfaces. The difference is the active ingredient, made visible. From there, three habits keep you honest. Judge any AI interviewing tool on the quality of its follow-ups and its completion rates, not the raw number of interviews. Always keep a human-moderated holdout, and report what it caught that the machine missed. And say plainly, wherever a machine did the interviewing, that it did, because a reader's trust in the finding depends on knowing who, or what, asked the questions.
Case studies
Microsoft Research, on the limits. Cuevas et al. (2025) ran an interview study with 399 participants, comparing two LLM-based chatbots against a fixed-question baseline. It is the largest peer-reviewed test of the method, and the most honest, because it found the chatbots met the generic quality marks but rarely captured people's specific motives or personal examples. The cautionary result and the state-of-the-art result are the same study.
Away and Outset, on scale. The luggage company Away, working through the vendor Outset, reports that one researcher ran seventy-five AI-moderated interviews overnight (Outset, 2025). It is a vendor account, not an independent one, so read the number as a claim. But it makes the shift concrete: work that would have taken a team weeks, done by one person in one night.
Conveo and Unilever, on the active ingredient. Conveo reports running two hundred interviews overnight for large consumer clients, and says most of the final insight came from the AI's follow-up questions, not the opening one (Conveo, 2026). Again a seller's own figure, but it states the active-ingredient claim plainly, from the people with the most data on it: the follow-up, not the first question, is where the understanding shows up.
User Intuition, on why real people still matter. In a self-published comparison, the vendor User Intuition ran the same study with real people and with synthetic respondents. The real participants pushed back on a flawed premise about a quarter of the time. The synthetic ones never did (User Intuition, 2026). It is a marketing study and should be treated as one, but it lands on a real point that connects to the wider debate about synthetic respondents. A real person interviewed by a machine can still tell you your question is wrong. A model imitating a person will not.
References
Chopra, F. and Haaland, I. (2026) Conducting qualitative interviews with AI. CESifo Working Paper No. 10666. Available at: https://doi.org/10.65864/ch6grwpore (Accessed: 14 August 2026).
Conveo (2026) AI-moderated research: insights and customer stories. Available at: https://conveo.ai/insights (Accessed: 14 August 2026).
Cuevas, A., Scurrell, J.V., Brown, E.M., Entenmann, J. and Daepp, M.I.G. (2025) 'Collecting qualitative data at scale with large language models: a case study', Proceedings of the ACM on Human-Computer Interaction, 9(2), pp. 1–27. Available at: https://doi.org/10.1145/3710947 (Accessed: 14 August 2026).
Gårdhus, T.P., Vitsakis, N., Frederiksen, F.L., Rogers, A. and Carlsen, H.B. (2026) AInterviewer: a platform for designing and conducting AI-led qualitative interviews. arXiv:2606.20588. Available at: https://arxiv.org/abs/2606.20588 (Accessed: 14 August 2026).
Geiecke, F. and Jaravel, X. (2024) Conversations at scale: robust AI-led interviews with a simple open-source platform. CEPR Discussion Paper DP19705 / SSRN 4974382. Available at: https://doi.org/10.2139/ssrn.4974382 (Accessed: 14 August 2026).
Höhne, J.K., Neuert, C.E. and Claassen, J. (2025) 'Effects of embodied interviewing agents on open narrative responses', International Journal of Market Research, 68(1), pp. 15–25. Available at: https://doi.org/10.1177/14707853251388213 (Accessed: 14 August 2026).
Horsfall, M., Hoogendoorn, A., Draisma, S., Eikelenboom, M. and Penninx, B. (2026) 'Interviewer and respondent sociodemographic characteristics, rapport, and their joint impact on data quality in the NESDA study', International Journal of Methods in Psychiatric Research, 35(2), e70090. Available at: https://doi.org/10.1002/mpr.70090 (Accessed: 14 August 2026).
Jacobsen, R.M., Cox, S., Griggio, C.F. and van Berkel, N. (2025) 'Chatbots for data collection in surveys: a comparison of four theory-based interview probes', Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, pp. 1–21. Available at: https://doi.org/10.1145/3706598.3714128 (Accessed: 14 August 2026).
Li, Z., Xiao, Z. and Li, T. (2026) Disclose with care: designing privacy controls in interview chatbots. arXiv:2602.01387. Available at: https://arxiv.org/abs/2602.01387 (Accessed: 14 August 2026).
Outset (2025) How Away used Outset to conduct 75 interviews overnight with a team of one. Available at: https://outset.ai/resources/stories (Accessed: 14 August 2026).
Rickwood, D. and Coleman-Rose, C. (2023) 'The effect of survey administration mode on youth mental health measures: social desirability bias and sensitive questions', Heliyon, 9(9), e20131. Available at: https://doi.org/10.1016/j.heliyon.2023.e20131 (Accessed: 14 August 2026).
User Intuition (2026) AI-moderated interviews and the Synthetic Mirage study. Available at: https://www.userintuition.ai (Accessed: 14 August 2026).
Williams, R.T. and Ingleby, E. (2025) 'The online survey in qualitative research: can AI act as a probing tool?', Frontiers in Research Metrics and Analytics, 10, 1519008. Available at: https://doi.org/10.3389/frma.2025.1519008 (Accessed: 14 August 2026).
Wuttke, A., Aßenmacher, M., Klamm, C., Lang, M.M., Würschinger, Q. and Kreuter, F. (2024) AI conversational interviewing: transforming surveys with LLMs as adaptive interviewers. arXiv:2410.01824. Available at: https://arxiv.org/abs/2410.01824 (Accessed: 14 August 2026).
Explore the idea
Let’s talk
Invisible forces shape your world — until you hire Latenta®
Contact