Synthetic Research Vendor Audits: The Demo Always Works

Article M2-09

A synthetic-sample vendor shows you a number that looks right, but only because they chose the test. Here is how to take that choice away from them.

In brief

A synthetic respondent is a language model prompted to answer survey questions as if it were a particular kind of person. A fast-growing market sells panels of them as a cheap replacement for human samples. The pitch usually arrives as a demo with an impressive accuracy figure. The problem is structural, not a matter of any one vendor's honesty. These models reproduce the average of a human sample while losing its spread, its subgroups and its variance. They are most accurate on the populations their training data covers best. A vendor who gets to choose the benchmark can therefore show you a real number that is also close to meaningless, because it was measured exactly where the model was always going to look good. The defence is an audit you control: you choose the held-out test, you aim it at the audiences the model has probably never seen, and you make the vendor disclose what it was built on. None of this is a named standard as of now. It is assembled from solid evidence about how these models fail and from professional codes only months old.

What the theory says

The theory

A synthetic respondent is not a person. It is a language model prompted to answer as though it were one, and vendors sell panels of these answers as a faster, cheaper stand-in for a recruited human sample. An audit has to establish where these models fail, because a good demo is designed to hide the failure.

The most corroborated finding in this young literature is that synthetic respondents reproduce the average of a human sample while breaking almost everything else about it. Large language models tested as replacements for survey data can match aggregate means while distorting the distribution of answers, the variance, and the relationships between variables, with the output shifting depending on the prompt and the model version (Bisbee et al., 2024). The AAPOR task force reached the same verdict in its 2026 review of AI in survey research, reporting that model-generated samples show less variation than human data and can distort variation and statistical relationships even when the averages look plausible (American Association for Public Opinion Research, 2026). This failure has its own full treatment as the flattening; here it is the mechanism that makes a flattering demo possible by construction. The number a buyer usually asks to see is an average, and the average is the only thing that remains accurate.

The second finding narrows down where the failure lives. Accuracy tracks the thickness of the training data. In a review of the use of these models in psychological research, models perform well on populations their training data covers heavily and degrade on populations it does not (Abdurahman et al., 2024). The point is put sharply by going outside the usual Western, educated, rich sample: testing synthetic participants against humans in global policy research, around seven in ten policy responses differed significantly from the real answers (Shrestha et al., 2024). So the failures follow a pattern. A synthetic panel looks most human where people like its training data are well represented online, and drifts furthest on the audiences that are thin on it. That shape tells the auditor where to point the test.

Put those two facts together and the audit is straightforward to design. It rests on three moves, none of them exotic.

The first move is choosing the test yourself. The accuracy figure a vendor reports is not a fixed property of their product. It is a function of which questions were asked and which people the answers were compared against, and the vendor picked both. A held-out test is simply a set of real human answers the vendor never saw, kept back so the model cannot have been tuned to them. If the vendor supplies the benchmark, they are grading their own work. If you supply it, from a real survey you already ran, the number starts to mean something. The mechanic here, training on one sample and testing on a separate held-out one, is ordinary validation and decades old. We cover the mechanic itself in holding out a sample to test a model. What is new is using it adversarially, as a buyer, against a vendor who wants to be tested in a way that makes them look good.

The second move is deliberately testing the model where it is weak. Because accuracy tracks training-data coverage, a test built from mainstream, well-covered questions will flatter almost any model. So you deliberately include the audiences the training data is thinnest on: people outside the dominant language and culture, specialist or low-incidence groups, anyone whose views are not all over the public internet. This is the opposite of what a demo does, since a demo shows the model answering a question it will answer correctly. Some audiences are so hard to reach that no held-out human sample exists to test against at all, which is its own problem and one we take up in audiences a model cannot reach.

The third move is asking what the model was built on. A vendor's synthetic panel is grounded on something: a base model, plus whatever data was used to steer it toward a market or a demographic. That grounding determines where the panel is accurate and where it collapses, and it is almost never disclosed. Asking for it is the simplest thing an auditor can ask for, and it has honest precedent outside market research. General machine-learning practice has developed documentation formats for exactly this, from the Data Cards proposed by Pushkarna et al. (2022) to the STANDING Together recommendations on dataset transparency in health data (Alderman et al., 2025). Neither is specific to survey research, and that gap is important, but they establish that describing what a model was trained on is a normal professional expectation.

The evidence that synthetic panels fail in these specific ways is strong and peer-reviewed. The audit procedure assembled from it is not. As far as this research could establish, there is no named, published audit protocol for synthetic-respondent vendors that a buyer could simply adopt. The closest thing is a set of guidelines from Sarstedt et al. (2024) for validating synthetic samples in marketing research, and those stop short of an adversarial, vendor-blind procedure. An independent psychometric audit by Lukauskas and Šarkauskaitė (2026) comes nearer: testing thirty-seven models against a held-out human sample, it runs several of these same moves and finds that models trained on synthetic answers lose their validity on the real people held back from them. It audits base models rather than handing a buyer a protocol to lift, but it shows the moves being run, not just proposed. The grounding-disclosure demand rests on professional codes that are only months old and on general machine-learning practice, none of it written for this exact case. So the argument here is not that the industry has a checklist you are failing to use. The pieces of one exist, scattered across solid evidence and thin standards, and assembling them is currently left to the buyer.

Auditing your data supplier is itself an old discipline. Buyers of human panels have long asked how a sample was recruited, screened and de-duplicated, drawing on a substantial literature on panel quality (Moss et al., 2023). The synthetic case needs a new audit because the old checks cannot be applied. There is no panel to inspect, no recruitment source to audit, no attention checks to review. There is a model and a prompt, and the only thing you can rely on is what it produces when you, not the vendor, choose the question.

Controversies

The main disagreement in the field is whether synthetic respondents can replace humans at all, and reported results vary widely. They range from near-parity to near-useless. The variation depends mostly on who chose the task.

On the encouraging side, a genuinely held-out test of synthetic purchase-intent ratings across dozens of real surveys reached roughly 90% of the reliability of the same humans answering twice. Maier et al. (2025) reported that result in a collaboration between the analytics firm PyMC Labs and Colgate-Palmolive. It is about the best result of its kind, and it is credible partly because the test was held out rather than hand-picked. On the discouraging side, an independent review by Lewis and Sauro (2026) counted findings across twelve papers since 2023 and found nine encouraging results and fourteen discouraging ones. Some studies replicated as rarely as one in five times. Sawtooth Software's Chapman (2026) found net approval figures wrong by up to thirty points and mean errors rising into the twenties. Two models from the same company disagreed on identical prompts, and both performed worst on underrepresented groups. An independent omnibus test by Verasight (2026), framed in mean absolute error, also warned against relying on synthetic samples.

These results do not really conflict once you see what varies. The favourable numbers tend to come from easy, well-covered tasks, often measured against the model's own distribution rather than against held-out human data. A vendor's choice of prompt can make the apparent match even better (Gui and Toubia, 2023). The unfavourable numbers come from harder tasks, held-out comparisons, or underrepresented groups. The score depends on who chose the test. That is also why a vendor's headline figure tells you so little by itself. Vendor claims are mostly high, from the 85 to 92% parity Synthetic Users (2026) reports, to the 98% figure Fairgen (2026) cites, to the 0.90 correlation Aaru (2026) reports against a third-party wealth study. Fairgen is explicit that its 98% is an internal plausibility gate rather than accuracy against humans. Almost none of these claims come with the four things that make a number research-grade: the benchmark used, which metric the percentage actually is, a margin of error, and the date it was last checked. A vendor page shows what is claimed, not that the claim is true.

There is a less obvious sign worth noting. Some vendors have stopped selling replacement. Roundtable AI (2026) has moved from offering synthetic humans to offering Proof of Human, a service for detecting AI-generated survey responses, which is nearly the opposite business. There is independent reason to take that change seriously. Westwood (2025), writing in PNAS, argues that large language models now threaten the integrity of online survey panels, because a model can produce believable answers in large numbers, which is exactly what a proof-of-human service exists to catch. Read charitably, the market is hedging its own strongest claim.

Limitations

The first limitation is the absence this whole article focuses on. As of now, there is no independent, standing service that runs vendor-blind holdout audits as a product, no certification body, no public leaderboard, and no accredited third-party auditor for synthetic-respondent vendors. A targeted search has so far returned nothing on point. The nearest thing is an independent benchmark from Chen et al. (2026), which runs one held-out protocol across several models and two separate human datasets and finds that none has yet beaten even the strongest non-model baseline at the individual level. It tests base models, not named commercial products, so the gap for vendors specifically remains. There is a useful public holdout dataset, Twin-2K-500, built by Toubia et al. (2025) from more than two thousand people answering over five hundred questions each, which gives an auditor real human answers to test against. But a dataset is not an audit. The procedure that turns it into one is what a buyer still has to run themselves.

The second limitation is about the method itself. A holdout audit is impossible or uninformative in three situations. When there is no real ground truth, because the audience or the product does not exist yet, there is nothing to hold out and compare against. When the audience is genuinely unreachable, the same problem appears in a different way, which is the territory of audiences a model cannot reach. And when the vendor is explicitly selling augmentation rather than replacement, a replacement-style holdout tests something the product never claimed to do. Fairgen (2026), for instance, positions its synthetic twins as directional rather than definitive, as a way to boost a real sample rather than replace it. Holding that to a replacement standard would be unfair. The audit applies squarely to the replacement claim, and you have to know which claim is on the table before you run it.

The third limitation is that the grounding-disclosure demand, the simplest thing the audit asks for, has no enforceable standard behind it. The ICC/ESOMAR International Code, in its 2025 fifth edition, states that a synthetic respondent is not a human participant and that its use must be disclosed, and it offers buyers guidance on asking about the augmentation method (International Chamber of Commerce and ESOMAR, 2025). AAPOR's 2026 framework adds transparency expectations. Both are recent, both state principles rather than mechanisms, and neither requires a vendor to describe their grounding data in the way a buyer needs. The demand is well-founded, but not yet backed by anything enforceable.

The fourth limitation is regulatory silence. ISO 20252, the main market-research quality standard, is silent on synthetic data in its 2019 edition, with a revision anticipated but not yet published (International Organization for Standardization, 2019). The EU AI Act covers synthetic content, but Article 50 concerns labelling AI-generated material, not whether a synthetic sample represents the population it claims to (European Union, 2024). On the question of validity, both are quiet, and that silence is itself the finding.

The fifth limitation is that the audit's own measurement may not reproduce. The procedure assumes that once you choose the test, the number you get back is a stable property of the vendor's model. A preregistered study by Zhu and Zhang (2026) puts that assumption in doubt for any audit run through a commercial model endpoint. Testing a black-box model as a measuring instrument on a shared endpoint, they sent byte-identical requests a day apart and got matching results in only 78 of 100 replays, against the 99 in 100 they had committed to in advance, while every engineering check passed: the request bodies matched exactly, and the model identifier and the other metadata the service returned were identical every time. Their task was arithmetic scoring rather than survey response, and they say plainly that a test whose options are well separated need not fail this severely, so the figure shows that the problem exists rather than measuring how large it is elsewhere. It is a preprint rather than peer-reviewed work, and one of its authors is affiliated with a commercial model vendor. The consequence for an auditor is cheap to act on: run the same held-out test twice in one sitting and check that the two results agree before treating either as the vendor's score.

The last limitation is originality. The core mechanic of the audit is classic validation, and this article invents no new statistic. What it assembles is the adversarial use of that mechanic, the coverage-aimed targeting, and the disclosure demand, pointed in a direction vendors would prefer you not look.

Open questions

Three unanswered questions determine how much a passed audit is worth.

Does passing a held-out audit certify anything beyond the test itself? If a vendor's synthetic panel matches your held-out human sample well, that doesn't guarantee accuracy on the next survey, a different question, or a real decision. No source has closed that gap yet. A passed audit tells you the model reproduced a specific set of answers it hadn't seen. That is necessary but not sufficient, so the sensible approach is to treat each new estimate as unproven until it is checked too.

Would the same audit, run across several named vendors at once, change the conclusion? Every independent finding so far is a review or a one-off study of a single model or product. No consortium has run one held-out protocol across competing vendors, which is the only way a buyer could compare vendors fairly. Until that exists, every comparison between vendors is really a comparison between studies that measured different things.

Does the coverage gap shrink as models improve, or is it built in? The audiences a model fails on today are the ones least represented in its training data. Whether better models close that gap, or whether it is intrinsic to how these models learn, decides whether the buyer's audit is a temporary protection or a permanent cost of using synthetic panels.

So what

The theory comes down to one uncomfortable fact and one defence. The fact is that a synthetic-sample vendor can show you a real, verifiable accuracy number that tells you almost nothing, because they chose where to measure it, and these models are built to look their best exactly there. The defence is not a better vendor or a stricter regulator, since neither exists yet. It is a small set of moves a buyer can run. You choose the held-out test yourself, aim it at the audiences the model has probably never seen, and make the vendor disclose what their model was grounded on and what their headline number actually measures. These methods can be run to find out what is true or to manufacture a number that flatters a decision already made, so the key distinction applies to every audience below. A study designed to agree with you is worthless as information.

For research practice

Treat every vendor accuracy figure as a claim to be reproduced, not a result to be accepted. The most useful habit is to keep a real survey you have already run, with real human answers, and never show it to the vendor. That becomes your held-out test. Ask the vendor to answer those exact questions with their synthetic panel, then compare. Aim at the parts of your audience the model is least likely to have seen, because matching on the easy majority doesn't prove much. When you report the result, say which benchmark you used, which metric the number is, and how uncertain it is, so nobody downstream mistakes a narrow test for a general guarantee. This is the same discipline that separates research from performance. The question is whether the test helps you make your decision, or whether it just gives you the answer you wanted.

For companies

Synthetic samples have real commercial appeal. They are fast, cheap, and available for audiences that are expensive to recruit. The mistake is to trust the accuracy figure on the slide. When a vendor reports 90-something percent parity, ask the four questions that turn a number into evidence: parity against which dataset, measured by which metric, with what margin of error, and last validated when. A vendor who can answer all four should be taken seriously. One who cannot has shown you a demo, and a demo is designed to work. Ask, too, what the model was grounded on, because that determines where it will fail without being noticed. Unnoticed failure is the expensive kind. A synthetic sample that is accurate for mainstream buyers and blind to a growing niche gives clean numbers and a wrong strategy. The error appears only when the product ships, after you have lost the early warning you were paying for.

For political parties

Synthetic panels are tempting for message and issue testing because they are fast and private. A party can check many variants without fielding anything. The risk is the same one the evidence keeps showing. The voters hardest for a model to imitate are the ones least represented in its training data: voters in a second language, older voters, voters outside the cultural mainstream. These are exactly the marginal groups an election can turn on.

A synthetic panel that says a message works has only told you it works with the well-covered majority. The rigorous and responsible approach is the same here. Use the synthetic panel to narrow a wide field of options quickly, then test the survivors on real people from the groups that actually matter. Do not trust the simulation to speak for them. A panel designed to agree with your existing beliefs is not useful information about the electorate. The vote is the held-out test you cannot choose.

For government and policy

Official statistics and public surveys do the most harm when a synthetic sample is wrong, because a wrong number leads to decisions that affect people. The people a model serves worst are often the same people public research is meant to help. There is still no standard a procurement officer can cite to require a synthetic vendor to prove representativeness. ISO 20252 remains silent on it, and the EU AI Act's synthetic-content rules govern labelling, not validity. So the buyer has to take responsibility and build the check into the contract. The defensible approach is to require a held-out audit against real human data the vendor never sees, aimed deliberately at the populations the programme is meant to serve, plus written disclosure of what the model was grounded on. Publish the result even when it is unflattering, because a negative finding about a tool the public is paying for is itself a public good. The decision test is strict. A synthetic sample chosen because it agrees with the answer a department already preferred does not save time. It makes a predetermined conclusion look like evidence.

How to use this

Run the audit in four steps. First, get a real held-out sample, a survey you have already fielded on humans and never shown the vendor. Second, make the vendor's synthetic panel answer those exact questions, and compare the full spread of answers, not just the averages. Run that comparison twice in the same sitting and check that the two results agree before you use either one (Zhu and Zhang, 2026). Third, aim the comparison at the audiences the model is least likely to know well, and treat a match on the easy majority as the start of the check, not the end. Fourth, ask for the two disclosures no vendor volunteers: what the model was grounded on, and the four facts behind the headline number.

A vendor who answers all of this is offering evidence. A vendor who offers a demo instead has not answered your question; they have only shown they can make the number look right when they pick the test. They can always do this, which is why you choose the test.

Case studies

Aaru and the wealth-study number (2026), the demo by construction. Aaru (2026) reports a 0.90 median correlation from recreating a third-party global wealth study with its synthetic panel. Taken at face value it is impressive. Looked at as an auditor would, it is a claim with the load-bearing details missing: no sample size, no significance, no held-out design, and no named party who computed it. The number may well be real. What it does not show is that the panel would reproduce answers on a test the vendor did not choose, which is the only test that would tell a buyer anything.

PyMC Labs and Colgate-Palmolive (2025), the honest counter-case. Maier et al. (2025), reaching roughly 90% of human test-retest reliability on purchase intent, did so on answers held out from the model rather than hand-picked, which is exactly why the number means something. It is narrated with a named client by a firm with an interest in the outcome, so read it as a demonstration of what an honest held-out audit looks like, not neutral proof that synthetic panels work everywhere. The lesson for a buyer is the shape of the test, not the size of the number.

References

Aaru (2026) Aaru: predictive intelligence. Available at: https://www.aaru.com/ (Accessed: 18 August 2026).

Abdurahman, S., Atari, M., Karimi-Malekabadi, F. et al. (2024) 'Perils and opportunities in using large language models in psychological research', PNAS Nexus, 3(7), pgae245. Available at: https://doi.org/10.1093/pnasnexus/pgae245 (Accessed: 18 August 2026).

Alderman, J.E., Palmer, J. et al. (2025) 'Tackling algorithmic bias and promoting transparency in health datasets: the STANDING Together consensus recommendations', The Lancet Digital Health, 7(1), pp. e64–e88. Available at: https://doi.org/10.1016/s2589-7500(24)00224-3 (Accessed: 18 August 2026).

American Association for Public Opinion Research (2026) Responsible AI Integration in Survey Research. Report of the AAPOR Task Force, 8 May 2026. Available at: https://aapor.org/announcements/task-force-on-responsible-ai-integration-in-survey-research-report/ (Accessed: 18 August 2026).

Bisbee, J., Clinton, J.D., Dorff, C., Kenkel, B. and Larson, J.M. (2024) 'Synthetic replacements for human survey data? The perils of large language models', Political Analysis, 32(4), pp. 401–416. Available at: https://doi.org/10.1017/pan.2024.5 (Accessed: 18 August 2026).

Chen, Z., Zhu, D. and Zheng, L.N. (2026) When synthetic users fail: a cross-domain benchmark of LLM-simulated human survey responses. arXiv:2607.26348. Available at: https://arxiv.org/abs/2607.26348 (Accessed: 19 August 2026).

European Union (2024) Regulation (EU) 2024/1689 of the European Parliament and of the Council (Artificial Intelligence Act), Article 50. Available at: https://eur-lex.europa.eu/eli/reg/2024/1689/oj (Accessed: 18 August 2026).

Fairgen (2026) Fairgen: boost your survey data. Available at: https://www.fairgen.ai/ (Accessed: 18 August 2026).

Gui, G. and Toubia, O. (2023) 'The challenge of using LLMs to simulate human behavior: a causal inference perspective', SSRN Electronic Journal. Available at: https://doi.org/10.2139/ssrn.4650172 (Accessed: 18 August 2026).

International Chamber of Commerce and ESOMAR (2025) ICC/ESOMAR International Code on Market, Opinion and Social Research and Data Analytics. 5th edn, 29 September 2025. Available at: https://iccwbo.org/news-publications/business-solutions/iccesomar-international-code-market-opinion-social-research-data-analytics/ (Accessed: 18 August 2026).

International Organization for Standardization (2019) ISO 20252:2019 Market, opinion and social research, including insights and data analytics: Vocabulary and service requirements. Available at: https://www.iso.org/standard/73671.html (Accessed: 18 August 2026).

Lewis, J.R. and Sauro, J. (2026) Can synthetic respondents replace human participants? A review of the evidence. MeasuringU, April 2026. Available at: https://measuringu.com/ (Accessed: 18 August 2026).

Lukauskas, M. and Šarkauskaitė, V. (2026) Plausible but not valid: a psychometric audit of LLMs as synthetic survey respondents. arXiv:2608.14606. Available at: https://arxiv.org/abs/2608.14606 (Accessed: 19 August 2026).

Maier, B.F., Aslak, U., Fiaschi, L., Rismal, N., Fletcher, K., Luhmann, C.C., Dow, R., Pappas, K. and Wiecki, T.V. (2025) LLMs reproduce human purchase intent via semantic similarity elicitation of Likert ratings. arXiv:2510.08338. Available at: https://arxiv.org/abs/2510.08338 (Accessed: 18 August 2026).

Moss, A.J. et al. (2023) 'Using market-research panels for behavioral science: an overview and tutorial', Advances in Methods and Practices in Psychological Science, 6(2). Available at: https://doi.org/10.1177/25152459221140388 (Accessed: 18 August 2026).

Pushkarna, M., Zaldivar, A. and Kjartansson, O. (2022) 'Data Cards: purposeful and transparent dataset documentation for responsible AI', in 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT '22), pp. 1776–1826. Available at: https://doi.org/10.1145/3531146.3533231 (Accessed: 18 August 2026).

Roundtable AI (2026) Proof of Human. Available at: https://roundtable.ai/ (Accessed: 18 August 2026).

Sarstedt, M., Adler, S.J., Rau, L. and Schmitt, B. (2024) 'Using large language models to generate silicon samples in consumer and marketing research: challenges, opportunities, and guidelines', Psychology & Marketing, 41(6), pp. 1254–1270. Available at: https://doi.org/10.1002/mar.21982 (Accessed: 18 August 2026).

Sawtooth Software / Chapman, C. (2026) Cautions on using large language models to simulate survey respondents. Sawtooth Software. Available at: https://sawtoothsoftware.com/ (Accessed: 18 August 2026).

Shrestha, P., Krpan, D., Koaik, F. et al. (2024) 'Beyond WEIRD: can synthetic survey participants substitute for humans in global policy research?', Behavioral Science & Policy, 10(2), pp. 26–45. Available at: https://doi.org/10.1177/23794607241311793 (Accessed: 18 August 2026).

Synthetic Users (2026) Synthetic Users: user research without the users. Available at: https://www.syntheticusers.com/ (Accessed: 18 August 2026).

Toubia, O., Gui, G., Peng, T., Merlau, D., Li, A. and Chen, H. (2025) 'Database report: Twin-2K-500: a data set for building digital twins of over 2,000 people based on their answers to over 500 questions', Marketing Science, 44(6), pp. 1446–1455. Available at: https://doi.org/10.1287/mksc.2025.0262 (Accessed: 18 August 2026).

Verasight (2026) Synthetic omnibus survey. Verasight, 13 January 2026. Available at: https://www.verasight.io/reports/synthetic-omnibus-survey (Accessed: 18 August 2026).

Westwood, S.J. (2025) 'The potential existential threat of large language models to online survey research', Proceedings of the National Academy of Sciences, 122(47). Available at: https://doi.org/10.1073/pnas.2518075122 (Accessed: 19 August 2026).

Zhu, H. and Zhang, J. (2026) Clean engineering, unstable measurement: a preregistered reliability failure of black-box LLM observers on shared endpoints. arXiv:2609.04198. Available at: https://arxiv.org/abs/2609.04198 (Accessed: 4 September 2026).

Explore the idea

Let’s talk

Invisible forces shape your world — until you hire Latenta®

Contact