Synthetic Expert Interviews: Interview the Expert You Could Never Book
Article M2-07
Reaching a surgeon, a chief financial officer or a rare-disease patient was always the hardest and most expensive job in research. Synthetic respondents promise to conjure them on demand, and the promise is loudest exactly where nobody has checked whether it holds.
In brief
Some audiences are close to impossible to survey: senior executives who will not take your call, specialist clinicians, patients with a one-in-ten-thousand condition. This is the case synthetic respondents are most promoted for, because the real alternative is so slow and expensive. The science behind the pitch is real but only partly related. Language models can out-predict human experts at forecasting which published result is true, and can approximate a general population's survey answers. Neither of those is the same as replacing an expert you cannot reach. The one test that would settle it, checking a synthetic surgeon against real surgeons the vendor did not choose, does not appear to have been published. Meanwhile the failure mode works in the wrong direction. These models make differences between groups more distinct rather than less distinct, so the clearest, most decision-ready synthetic answer is the one most likely to be invented.
What the theory says
The theory
Some people are difficult to survey. A chief financial officer will not spend forty minutes answering your questionnaire. A transplant surgeon is part of a small population, spread across hospitals, and costly to recruit for an hour of their time. A patient with a rare disease may be one of a few thousand people in the country, and reaching a representative few of them can take a specialist agency months. Market research has always had methods for this, and they are not new or special. Delphi expert elicitation runs several rounds of structured questions past a panel of specialists, showing the group's answers back to them until the estimates stabilize, and it is still in active use for questions no dataset can answer (Sohst, Acostamadiedo and Tjaden, 2023). Respondent-driven sampling reaches hidden or low-incidence groups by letting members recruit each other, with a statistical correction for the way the chain grows (Jayaweera et al., 2023). Both are slow, both are costly, and both involve real people. That is the standard the new pitch argues against.
The change since 2022 is that a language model offers to avoid recruiting entirely. A group of vendors that did not exist three years ago now sells synthetic audiences designed to represent exactly these hard-to-reach groups, among them Evidenza, Aaru, Artificial Societies and Qualtrics Edge Audiences. The pitch is most persuasive here, for a simple reason. When the real audience costs a lot and takes months, a simulation that answers in seconds looks like an obvious choice. The unreachable-expert case is the strongest commercial story in the whole synthetic-respondent field, which is also why it deserves the closest examination.
This is a different problem from copying someone you can actually interview. The interview-grounded digital twin of a named person has a measured accuracy ceiling precisely because a real transcript anchors it and can be scored against that person's own answers. Here, by definition, there is no interview and no real member of the audience to anchor against, which is why the validation that gives a twin its number is the one this case cannot get.
It has a real scientific result behind it, and that is what makes it difficult to dismiss. Luo et al. (2024) built a test called BrainBench that asked whether a study's real result could be told apart from a plausible fake one, and found that language models did this better than human neuroscientists. Ashokkumar et al. (2026), across seventy survey experiments, found that a model's predictions of the outcomes correlated about 0.85 with what the experiments actually found. Those are strong results, and a vendor can reasonably point to them. But they answer a related question, and the difference matters. Both test whether a model can predict a published or aggregate result. Neither tests whether a model can be a particular expert respondent, giving the private, unpublished judgement that a real CFO or surgeon would give to a question nobody has answered in print. Predicting the field's result and acting as a member of the field are different tasks. The evidence is strong on the first and almost absent on the second.
Controversies
The argument in the field is not whether these models do anything useful. It is whether they do the job they are sold for. Academic teams that build synthetic panels tend to frame them as a complement to human experts rather than a replacement. Boone et al. (2025), validating persona-driven models against instruments that measure children's trust, position the tool as something to run alongside real respondents, not instead of them. The vendor framing is the opposite. The commercial promise is that you no longer need the hard-to-reach human at all. That gap, between complement and replace, is the controversy.
The evidence that should worry a buyer is about how these models fail, not whether they work on average. Chen, Zhu and Zheng (2026) ran a benchmark of model-simulated survey responses across several domains and found two things that are directly relevant to the expert case. First, no model beat the strongest simple baseline at the level of the individual respondent. Second, and more revealing, the models over-determined demographics. They inflated the gaps between segments, making the difference between one professional group and another two to four times larger than it is in a real population. Call this over-determination. It is a distinct failure from the more familiar one where a model drifts toward the average and loses a group's internal variety, which is the variance collapse problem covered elsewhere in this corpus. Over-determination works in the opposite direction. The model does not blur the segments together. It separates them too much. That has an uncomfortable consequence. The crisper a synthetic answer looks, the more cleanly it separates the CFO from the CMO, the more likely that separation was manufactured rather than found.
A second mechanism directly challenges the expert claim. Hu, Rostami and Thomason (2026) found that giving a model an expert persona improved how well its answers matched what an expert sounds like, while making the answers less accurate. Telling a model to be a cardiologist makes it talk more like a cardiologist and reason slightly worse. That is the specific trap of this whole category. Expert language and expert judgement are not the same, and the model is better at producing the first than the second. A synthetic expert can be good at using expert language and sounding confident, but often wrong in a way that is hard to notice because it sounds correct.
Limitations
Two limits apply to all of this. The first is that the strongest evidence for the method does not directly address the target. The BrainBench result is about predicting which published finding is real, which is matching patterns in text the model has already seen. A real expert respondent's value is the opposite: the judgement they carry that is not in any published text, the insight that comes from practical experience. Ashokkumar's result is about a general population meant to mirror the country, and even there the model tends to overestimate how big an effect is and has blind spots on sensitive topics. Both results are the best possible for the easy case. The hard case, a specific expert who cannot be reached, is compared to that ceiling and has no direct test at all.
The second limit is in how accuracy gets reported. Vendors publish a single aggregate number: an average correlation or an overall similarity score, often between 0.8 and 0.9, sometimes rendered as a headline percentage. Artificial Societies, for instance, reports 86% distribution accuracy, benchmarked against general-population university surveys rather than any expert or low-incidence audience (Artificial Societies, 2026). That kind of number is real, but it is not the right one. An average across a whole audience can look excellent while the model is badly wrong in each specific part, because the errors cancel out in the average. The metric that would actually reveal a failure in expert areas, the accuracy of the answer within each specific segment or the shape of the full distribution rather than its average, is the one not reported. And the vendor chooses which benchmark to report against, so the party being measured also chooses the test.
The clearest sign that even the vendors know this comes from what they build. Qualtrics, selling a synthetic business-to-business audience, grounds the model in real panel data through a partnership with a panel firm rather than trusting the model alone (Qualtrics, 2025). That is an indirect admission that the hard-to-reach case needs a base of real human data underneath it, which is the thing the pure-simulation pitch says you can skip. The nearest independent measure outside the vendors points the same way. Toubia et al. (2025) built an open benchmark of digital twins of over two thousand real people and measured how faithfully models reproduce them. It is a general-population test, not an expert one, and it is a measure rather than a validation of the hard case, but it exists in the open where the vendor numbers do not.
No standards body fills that gap either. The 2025 ICC/ESOMAR Code and ESOMAR's buyer guidance both govern disclosure, requiring you to say the respondents were synthetic so no one is misled into thinking real people were surveyed, but neither certifies that a synthetic answer is accurate (ICC/ESOMAR, 2025; ESOMAR, 2025). There is no certification to pass, and no standards body is offering one.
Open questions
Three questions the field has not answered decide whether a synthetic expert is worth trusting at all.
Does the expert-prediction win transfer to expert-respondent simulation? Predicting a published result and giving a working expert's private, unpublished judgement are different tasks, and nothing yet bridges them. The transfer might hold or it might fail completely, and nobody can currently say, so a vendor claiming it is describing a hope rather than a result.
Would a synthetic expert survive a held-out test against real experts the vendor did not pick? As far as this research could establish, no peer-reviewed study validates a synthetic respondent against real expert or low-incidence ground truth, and every benchmark that does exist is run against an audience reachable enough to measure, which is precisely not the audience the method is sold to replace. The validation is missing exactly where willingness to pay is highest. Until an adversarial, held-out check of the kind set out in the audit procedure is published, the simulation must prove it works, not the buyer.
Is over-determination a fixable artefact or intrinsic to persona-prompting? Models draw the gaps between segments two to four times too sharply (Chen, Zhu and Zheng, 2026), and giving a model an expert persona makes it sound more expert while reasoning slightly worse (Hu, Rostami and Thomason, 2026). If those are tuning problems, better engineering fixes them. If they are intrinsic to how a language model plays a role, then the more polished and useful a synthetic expert answer seems, the less it can be trusted, and no amount of prompting makes it safe to read straight.
So what
The claim of synthetic expertise is most persuasive where it is least verified, and these models' failures mean the most confident answer is the least trustworthy. That combination makes the method risky, not just new. A tool that produced obvious nonsense would be harmless, because you would discard it. This one produces clear, expert-sounding, ready-to-use answers about people no one in the room can contact to challenge. The danger is not that you will lie with it. The danger is that you will believe it, because it tells you something specific and plausible about an audience you could not otherwise access. So the first move in every use below is the same. Treat a synthetic answer about an unreachable audience as a claim to verify, not as evidence, and admit that the verification has usually not been done.
For research practice
You are accountable for the number you hand over, whether a real person or a model produced it. So the requirement is the same as for any other tool: what was it validated against, and by whom. For a synthetic hard-to-reach audience the answer is almost always that it was validated against a reachable proxy, and that the vendor chose the comparison. Ask for the per-segment or full-distribution match rather than the headline correlation, and if only the average is on offer, treat that as a limited result. The usable approach is the one supported by academic research. Use the synthetic panel to quickly reduce a large set of questions or hypotheses, then spend your real budget testing the survivors on however few real experts you can actually reach, even if it is five. That sequence uses the model usefully without making it the final decision-maker. A smooth, neat result from a synthetic expert is the least reliable result possible, because a fluent model gives a confident answer whether the answer is correct or incorrect.
For companies
The unreachable business buyer is the most valuable audience a synthetic vendor can offer and the least checked, which is not accidental. It is expensive to reach in real life, which is exactly why a simulation is worth most there, and also why validating that simulation is hardest there. When a vendor reports that its synthetic executives match real ones at 88%, two questions matter most: what exactly was matched, and who picked the audience it was matched against. The key question is whether the study helps your decision. Does this study serve the decision you are making, or the conclusion someone already wanted? A synthetic panel is especially likely to produce the second kind of study, because a model nudged with the right persona tends to agree in the register you primed it with. A panel that agrees with your plan is not useful information. It just reflects what you already wanted. A flattering simulation removes the early warning of problems that you were paying for research to get.
For political parties
This is where a manufactured answer causes the most harm, and the way it fails makes that harm worse. A campaign wants a clear understanding of a hard-to-reach part of the electorate: a group that will not answer polls, a narrow band of persuadable voters. That is exactly the case where a model over-determines the segment and draws the difference two to four times too sharply (Chen, Zhu and Zheng, 2026). The simulation will give you a clear, confident account of how that group thinks, and that clarity is part of the problem. A synthetic panel designed to tell you your message works is the same thing as a push poll: a tool built to produce the answer you wanted rather than to find out what is true. The decision test is the safeguard. If the aim is to understand the electorate, the rigorous approach and the honest approach are the same, which is to use the model to narrow a large field of messages fast and then test the survivors on real people from the groups that matter. If the aim is a number that justifies a decision already made, the model will readily do so.
For government and policy
Government research reaches for exactly the audiences that are hardest to simulate honestly: rare-disease patients, minority-language communities, small professional bodies, the people official statistics most need and least reach. These are the highest-stakes uses and the least checkable, because there is no large real sample to validate against. That is the case for holding the tool to its standard rather than avoiding it. If a synthetic panel stands in for a population that a policy will affect, the disclosure rule is the floor, not the ceiling. The 2025 ICC/ESOMAR Code requires you to state that the respondents were synthetic and not imply that real people were consulted (ICC/ESOMAR, 2025). Above that floor, the responsible practice is to validate against whatever real sample can be found, however small, and to publish the shortfall when the model falls short rather than only the best case when it does not. A synthetic tail is not evidence about the real tail.
How to use this
Before you trust a synthetic answer about an audience you cannot reach, ask three things. First, was the model ever checked against real people of this exact kind, or only against a reachable stand-in the vendor found convenient? Second, was the reported accuracy a match on the full distribution and the individual segments, or a single average correlation that can hide large errors cell by cell? Third, who chose the test, and could you run your own on questions the vendor has never seen? If the answers are a reachable proxy, an average, and the vendor, you have a hypothesis generator, not a source of evidence, and you should price it accordingly. The full procedure for running that independent check, the held-out validation against real experts of the same kind, is the audit we set out for the general case. The harder an audience is to reach, the more the burden of proof should sit on the simulation, not on you.
Case studies
Evidenza. The sharpest version of the unreachable-expert pitch comes from Evidenza, which sells AI copies of hard-to-reach business buyers and names the exact roles that are nearly impossible to book, from chief executives to chief information and chief financial officers. Figures such as 88% accuracy, a 95% same-conclusions rate against an EY study, and a 0.81 correlation with Salesforce data appear in the company's marketing and in secondary coverage, while the FAQ itself lists no numbers (Evidenza, 2026). Every reported benchmark is an aggregate correlation against a reachable audience, on a test the vendor selected. The metric that would expose an expert-domain failure is the one not shown, and the party being measured picked the measure.
Qualtrics Edge Audiences. The incumbent's design is the tell. Rather than trust a model alone for its synthetic business audience, Qualtrics grounds it in real panel data through a partnership with a panel firm (Qualtrics, 2025). Grounding the hard case in real human data is a candid admission that the pure simulation is not enough where it counts.
Aaru. Aaru's public proof points are a New York primary predicted close to the result and a wealth survey correlated with an EY benchmark (Aaru, 2026). Both are reachable, measurable outcomes, and neither is an unreachable expert audience. An investment from a large consultancy is capital, not validation. The strongest evidence on offer is for the case the method finds easy, not the case it is sold for.
References
Aaru (2026) Behavior simulation at the scale of the real world. Available at: https://aaru.com/ (Accessed: 18 August 2026).
Artificial Societies (2026) Audience simulation for concept, ad and video testing. Available at: https://societies.io/ (Accessed: 18 August 2026).
Ashokkumar, A. et al. (2026) 'Large language models can predict the results of social science experiments', Nature, 656(8126), pp. 115–122. Available at: https://doi.org/10.1038/s41586-026-10742-x (Accessed: 18 August 2026).
Boone, E. et al. (2025) 'Synthetic Validation of Pediatric Trust Instruments using Persona-Driven Large Language Models', medRxiv [Preprint]. Available at: https://doi.org/10.1101/2025.11.25.25340922 (Accessed: 18 August 2026).
Chen, Z., Zhu, D. and Zheng, L.N. (2026) 'When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses', arXiv [Preprint]. Available at: https://doi.org/10.48550/arXiv.2607.26348 (Accessed: 18 August 2026).
ESOMAR (2025) ESOMAR standards and guidelines. Available at: https://standards.esomar.org/about (Accessed: 18 August 2026).
Evidenza (2026) Survey AI copies of your customers, even the hardest-to-reach B2B buyers. Available at: https://www.evidenza.ai/ (Accessed: 18 August 2026); FAQ: https://www.evidenza.ai/faqs (Accessed: 18 August 2026).
Hu, Z., Rostami, M. and Thomason, J. (2026) 'Expert Personas Improve LLM Alignment but Damage Accuracy: Bootstrapping Intent-Based Persona Routing with PRISM', arXiv [Preprint]. Available at: https://doi.org/10.48550/arXiv.2603.18507 (Accessed: 18 August 2026).
ICC/ESOMAR (2025) ICC/ESOMAR International Code on Market, Opinion and Social Research and Data Analytics. 5th edn. Paris/Amsterdam: ICC/ESOMAR. Available at: https://iccwbo.org/news-publications/business-solutions/iccesomar-international-code-market-opinion-social-research-data-analytics/ (Accessed: 18 August 2026).
Jayaweera, R.T. et al. (2023) 'Respondent-Driven Sampling for Estimation of the Cumulative Lifetime Incidence of Abortion in Soweto, Johannesburg, South Africa: A Methodological Assessment', American Journal of Epidemiology, 192(7), pp. 1081–1092. Available at: https://doi.org/10.1093/aje/kwad074 (Accessed: 18 August 2026).
Luo, X. et al. (2024) 'Large language models surpass human experts in predicting neuroscience results', Nature Human Behaviour, 9(2), pp. 305–315. Available at: https://doi.org/10.1038/s41562-024-02046-9 (Accessed: 18 August 2026).
Qualtrics (2025) Edge Audiences: human + synthetic research. Available at: https://www.qualtrics.com/strategy/audiences/ (Accessed: 18 August 2026).
Sohst, R.R., Acostamadiedo, J.M. and Tjaden, J. (2023) 'Reducing uncertainty in Delphi surveys: A case study on immigration to the EU', Demographic Research, 49, pp. 983–1020. Available at: https://doi.org/10.4054/demres.2023.49.36 (Accessed: 18 August 2026).
Toubia, O. et al. (2025) 'Database Report: Twin-2K-500: A Data Set for Building Digital Twins of over 2,000 People', Marketing Science, 44(6), pp. 1446–1455. Available at: https://doi.org/10.1287/mksc.2025.0262 (Accessed: 18 August 2026).
Explore the idea
Let’s talk
Invisible forces shape your world — until you hire Latenta®
Contact