Synthetic Sample Validation: The Mean Matches. That Is the Trap.
Article M2-03
A synthetic sample can hit the average of a real survey and still be wrong in five distinct ways, each needing its own fix or having none at all. The one number that matches is what hides all five.
In brief
Synthetic respondents are language models asked to stand in for the people a survey would normally recruit. They are cheap, fast, and convincing on the one thing everyone checks: the average. A model can reproduce the headline number of a real study and still get almost everything else wrong. The problem has a name worth keeping, the flattening, and it is not one fault but at least five that a matched mean conceals. The spread of opinion collapses toward the middle. The extremes vanish, so the enthusiastic minority and the hard no disappear. The model agrees with whoever is prompting it. It answers like a well-read Westerner no matter who it is meant to be. And it knows nothing after its training cutoff. Two of these failures come straight from the fine-tuning that makes the model useful, which is why the flattening is so hard to remove. The finding that the average stays accurate while the overall distribution becomes inaccurate is well evidenced now. The fixes are not: no published method rebuilds the full distribution, and the best industry evidence makes the average more precise but does not restore any missing response patterns.
What the theory says
The theory
In this article a synthetic respondent is a large language model prompted to answer a survey as if it were a person, or a demographic slice of people. The appeal is clear. Recruiting a thousand real respondents is slow and expensive, and a model can produce a thousand answers in minutes. If those answers matched what real people would say, this would be one of the largest cost reductions market research has ever seen.
The trouble is what a match means. The number almost everyone checks first is the average: what share of the sample picked option B, how the group leans on a five-point scale. On that test synthetic samples do reasonably well, and that early success is what has attracted investment. Moon and colleagues (2026) ran the most careful version of this check to date, testing several models on their ability to reproduce not just the average of a human survey but its whole distribution. The models reproduced condition-level patterns tolerably and failed to reproduce the distributional structure, the way real answers spread out across the options. The result that should strongly discourage a buyer is this: on their benchmark, no model beat a baseline that ignored the survey conditions entirely. A synthetic sample matched to the mean was, on the measure that matters, further from the humans than a crude pooled guess. Bisbee and colleagues (2024) found the same failure in political survey data, where model-generated opinions came out less varied than real ones and shifted unpredictably with the prompt and the date.
This is the flattening, and it is not one defect. [M2-01] and [M2-02] both hand the distribution failure to this article. A matched average hides the problems. Under it are at least five distinct failures. The failures differ in how they can be fixed, and several cannot be fixed at all. They all show up as the same symptom: a spread that is too narrow and missing extreme values. So a buyer who checks only the mean cannot tell them apart, and a vendor who has fixed one can honestly call the sample improved while the rest remain. A mitigation aimed at one failure does nothing for the others. Nor can you buy your way out by generating more respondents. A large biased sample gives false precision, not accuracy, the older lesson Yang and colleagues (2024) documented for real survey data. A hundred thousand synthetic answers from the same flattened model give you a narrower range around the same wrong distribution.
The five failures do not each have a separate cause. They trace to a smaller set, and the largest cause is the step that makes these models usable at all. A raw language model is trained to predict text; a helpful assistant is then produced from it by reinforcement learning from human feedback, usually shortened to RLHF, in which people rate model answers and the model is tuned toward the ones people prefer. Kirk and colleagues (2024) measured what that step does to the range of a model's output and found it significantly reduces diversity across several ways of counting it. The same tuning that makes a model generalise well makes it say a narrower set of things, and it also teaches the model to agree. So two of the five failures, the collapse of the spread and the habit of telling you what you want to hear, come directly from the one process that made the model useful. That shared cause is why the flattening is hard to fix. You cannot remove the failures without changing the step that makes the tool work, and no amount of better prompting removes a cost built in that deeply.
The first failure is variance collapse. The spread of opinion contracts toward the centre, so the model reproduces the average person and loses the disagreement. [M0-03] owns variance collapse as a validity failure, and its one partial fix belongs there too. What this article owns is why the spread narrows in the first place, and the answer is the RLHF engine above.
The second failure is the one a buyer notices first: the tails go missing. When the spread collapses it does not shrink evenly. The extreme positions thin out fastest, so the enthusiastic minority and the hard no, the very answers that matter for a new product or a divisive issue, are the first to vanish. This is the same narrowing Moon and Bisbee measured, measured from the edge of the distribution rather than its middle, so it shares variance collapse's partial fix and has no other.
The third failure is sycophancy, the model's tendency to tell the prompter what the prompter seems to want. This is the second failure the RLHF engine produces. Sharma and colleagues (2024) tested several leading AI assistants across a range of open-ended tasks, found the behaviour in all of them, and traced it to the same preference data that narrows the spread, because that data rewards answers matching the user's stated view. For a survey this is close to disqualifying, because a synthetic respondent that leans toward the questioner's apparent position is not independent. Ask a leading question of a sycophantic model and it complies, which is the failure mode a real survey's methodology tries to prevent. Its fix is different from the first two: careful, non-leading prompting rather than held-out data.
The fourth failure shows up when you do not tell the model who to be, and its cause is different. Not the tuning but the composition of the training text. When not given a specific role, these models answer like a particular kind of person: Western, educated, comfortable, from a rich democracy. Researchers call that profile WEIRD, for Western, Educated, Industrialised, Rich and Democratic. Atari and colleagues (2023) measured how closely model responses track real human data across many countries and found the resemblance highest for WEIRD populations and decreasing sharply for everyone else. Santurkar and colleagues (2023) found the same skew inside a single country, reporting that the gap between model opinions and particular US demographic groups was about as wide as the gap between Democrats and Republicans. A synthetic sample can be told to role-play a group it does not know, and it will answer anyway, smoothly and incorrectly. There is no reliable fix.
Frozen time is the fifth failure, and its cause is the simplest: the training cutoff. A model knows nothing that happened after its training data was collected, so an attitude that formed in response to a recent event, a new product, a scandal, a price shock, is outside what it can represent. There is no published study of a synthetic survey sample being wrong specifically because the opinion post-dated the cutoff. The support is a mechanism argument by analogy: Dai, Teehan and Ren (2024) showed that models degrade in a structured way when reasoning about events past their training horizon. Treat frozen time as a mechanism you can reason about, not a measured survey failure, because that is what the evidence currently shows. It has no measured fix.
Controversies
The live argument is not whether the flattening happens. It is whether it can be fixed now, and here the honest answer is that describing a distribution and reproducing one are different achievements. Meister, Guestrin and Hashimoto (2024) made that distinction clearly: a model can more accurately describe what the distribution of opinion looks like than actually generate a sample that has that distribution, and they call reliable simulation an open problem rather than a solved one. That difference is easy to overlook in a demo, because a model that can describe the spread looks like a model that can produce it.
Against that sit a growing set of mitigation claims. Huang, Li and Shao (2025) report that aligning a model against the way distributions shift helps it simulate survey response distributions more faithfully. Jia and colleagues (2026) approach the issue from the opposite direction, mapping the conditions under which digital personas can approximate a human survey at all. Both suggest a useful approach. Neither has been independently replicated, and neither meets the important standard, a full distribution that survives a genuinely held-out test. The cautious interpretation is that these are demonstrations, not confirmed results. When a vendor shows a repaired distribution, ask whether the repair was checked against a held-out sample, because matching the data it was trained on proves very little.
The strongest near-independent industry evidence supports the same conclusion. Even the best result the field has, Fairgen's Boost, reduces uncertainty without restoring the lost variation or the missing extremes. The case study below carries the numbers.
Limitations
The main limitation is that no independent standard verifies whether a synthetic sample has the correct distribution. The governance that does exist regulates something else. The ICC/ESOMAR International Code, updated in 2025, clearly distinguishes a real person from a synthetic persona and sets rules for labelling and human oversight. Those are worth having. They say nothing about whether the distribution a synthetic sample produces matches reality, because that is not what a code of conduct is for. So a buyer cannot rely on a standards body to answer the distribution question. There is no ISO standard that confirms the distribution tails are accurate.
The WEIRD skew has a measured cost to a specific, market-relevant subgroup, though here the evidence narrows to a single study. Kim and colleagues (2025) report that language models systematically misrepresent American climate opinions, getting a concrete, single-topic distribution wrong in a way that would mislead anyone using a synthetic sample to read that issue. As far as this research could establish, that is the clearest named instance of the skew affecting a real subject, and it rests on one preprint. The population-level pattern is well supported by Atari and Santurkar. The specific worked example is one paper, and a reader deciding how much to trust it should know that.
One caveat makes the comparison fair. Real surveys have their own well-known distortions, and the claim here is not that human data is a flawless standard. It is narrower: whatever errors a real sample has, a mean-matched synthetic sample adds a separate, measurable loss of distribution accuracy.
Open questions
Three questions the field has not settled decide how much weight a synthetic sample can carry.
Is the flattening a limit of today's models, or of the whole approach? Two of the five failures come from the RLHF step that makes a model useful in the first place, so it is not obvious that a bigger or better-trained model would flatten any less. The answer tells a buyer whether to treat synthetic samples as a young technology worth waiting on, or as a tool that is permanently for exploration only.
Can any mitigation rebuild the full distribution and survive an independent test? The fixes that exist today work only in pieces, and none has been shown to hold up on data it was not built against. That answer is the line between a synthetic sample staying a cheap first screen and becoming something you could genuinely measure with.
Does frozen time actually bite, and how hard? No one has yet measured a real case of a synthetic sample getting an answer wrong because the opinion formed after the training cutoff. It is the one failure with no empirical grounding at all, so a buyer cannot size the risk, and it is most likely to strike exactly the recent-event, fast-moving questions that research is often hired to answer. That no case has been found means the evidence is missing, not that the failure does not happen.
None of the five failures is solved, and that is the honest state of the field today.
So what
The main point is the same for every audience, and it is simpler than the five failures that cause it. A synthetic sample reliably reproduces the average and unreliably reproduces everything else, and most decisions depend on the second half. Many decisions depend on the spread of opinion rather than its centre: how divided a market is, which minority is about to matter, where the strongest opinion is. On all of those a synthetic sample is least accurate in the areas you need most. There is also an ethical issue here before a practical one. Because these models lean toward the answer the prompter seems to want, a synthetic sample can be steered, by loaded prompting or by picking the persona that returns the convenient result, into agreeing with the conclusion you already preferred. The defence is the same as against a rigged real survey. Ask whether the exercise serves the decision being made or the answer the sponsor wanted, because a sample built to agree with you is useless as information. Such a sample tells you about the prompt, not the population.
For research practice
The main risk here is false agreement. If a synthetic sample comes back showing agreement, that is the least trustworthy result it can give you, because a model that smooths differences can show agreement even when the real population disagrees. So the practical rule is: treat any synthetic result as a guess about the average, not about how spread out the answers are, and check it against real respondents before you use it to make a decision that depends on how spread out the answers are. Synthetic samples are good for a quick, low-cost estimate of the average and for ranking clear choices before you pay for real surveys. They cannot measure disagreement, estimate the size of a minority, or understand a group that differs from the model's usual profile, because those all depend on the variation that the model removes.
For companies
The sales message is that synthetic samples replace human panels: you skip the panel and get the same result for less money. The honest version is more limited but still useful. Synthetic samples can improve the minimum quality of many small studies that currently get no research, giving a rough average result where the alternative was a guess. The mistake is to use them for the extremes of your market, because that is exactly where they fail. New-product and concept work depends on finding the enthusiastic minority and the strong negative responses, and those disappear first when the range of responses narrows. When a vendor says its synthetic sample matches human data at a high rate, ask the question that distinguishes the sales pitch from the actual product: matches on what? [M2-09] sets out how to audit that claim. A match on the average is what the evidence supports and what Fairgen's near-independent study delivered, a narrower interval by extrapolation. A match on the full distribution, including the extremes and the disagreements, is what no one has yet shown holds up under an independent test. If you buy the first, you have a useful, limited tool. If you believe you bought the second, you ship a product to a market you never measured.
For political parties
Wording and framing move polling numbers, and a party that can test many message variants quickly and privately has a real edge. Synthetic samples make that volume affordable. The warning is stronger here than anywhere else, for two reasons that add up. First, the voters hardest for a fluent model to impersonate are the ones with less formal education, older voters, and voters answering in a second language, which is a large and often decisive part of any real electorate. A synthetic sample that says a message lands has really only told you it lands with a well-read analyst. Second, and this is the ethical rule that must come first, not last, these models tilt toward the answer the prompter wants, so a synthetic panel can easily be used to confirm the strategy you had already chosen. The rigorous use and the responsible use are the same action: narrow a wide field of wordings fast with the model, then test the survivors on real people from the groups you actually need, and never let a simulation that agrees with you replace a coalition that might not. A synthetic panel that only confirms the chosen strategy tests nothing, and the election depends on the very voters it cannot impersonate.
For government and policy
Official statistics and public consultation are most harmful when the sample is unrepresentative, because the people an unrepresentative instrument serves worst are often the ones the exercise exists to reach. A synthetic sample imports the WEIRD skew directly into that setting, representing the citizens already best represented most accurately and everyone else least accurately, which defeats the purpose of a public statistic. The defensible use is as an initial tool that expands coverage at low cost. Pair it with real data for any number that guides a decision or gets published. Keep an internal record of which findings came from a model and which from people. Two lines should be non-negotiable. A synthetic sample never supplies a headline official figure on its own, and a synthetic result is never presented without saying it is synthetic, because the ICC/ESOMAR code and its peers are explicit that a persona is not a person and the label must accompany the number. Precision from extrapolation is not the same as having asked the people directly.
How to use this
Treat the matched average as the start of the check, not the end. Before you trust a synthetic sample, ask four questions. Does the decision depend on the spread of opinion or only its centre, because if it is the spread, the model is least accurate in that area. Could the result have been steered, by a leading prompt or a hand-picked persona, toward the answer someone wanted, because sycophancy makes that easy and hard to see. Does the sample need to speak for a group far from the model's default profile, because the further out you go the less it knows. And is any part of the opinion tied to something recent, because the model has no information after its training cutoff. If the honest answer to any of these is yes, the synthetic sample is a first pass and not the measurement, and the real respondents it was meant to replace are the ones you still need. The technology now provides a cheap average where none existed before. It has not made the disagreeing, unusual, recently-changed, real respondent optional. A matched average that makes you think otherwise is where the problems begin.
Case studies
Verian Group (2026). The research agency formerly known as Kantar Public published a critical assessment of synthetic sample in social research, useful precisely because it comes from inside the industry rather than from a lab. Their synthesis is blunt about the shape problem: synthetic samples flatten variance, misportray identity groups, and manufacture false precision, and they note that a synthetic exercise can look as confident as a real survey with several times its effective sample size. It is a position paper rather than their own experiment, and worth reading as a named agency arguing against a product it could otherwise be selling.
Fairgen Boost with Google Japan and Macromill (2025). The field's one near-third-party validation, presented through ESOMAR, reported a mean confidence-interval improvement of roughly a quarter across fifty-two data cuts on real market-research data. It is the best evidence the industry has, and it still does not claim shape recovery. The gains come from extrapolating larger subgroups onto smaller ones, need a floor of a few hundred real respondents, fade to almost nothing at the smallest samples, and carry the vendor's own warning that the method cannot surface a finding that was not already present. Even the strongest case narrows an interval; it does not rebuild a distribution.
Qualtrics Edge (2025). Launched in March 2025 and built on a very large pool of third-party respondents, Edge is the scale-and-momentum case rather than an evidence case. Its published claims are about cost, speed and point accuracy, all self-reported, with no independent measurement of variance or tail fidelity. It belongs here as a reminder that commercial momentum and distributional evidence are moving at very different speeds.
References
Atari, M., Xue, M.J., Park, P.S., Blasí, D.E. and Henrich, J. (2023) 'Which humans?', PsyArXiv preprint. Available at: https://doi.org/10.31234/osf.io/5b26t (Accessed: 18 August 2026).
Bisbee, J., Clinton, J.D., Dorff, C., Kenkel, B. and Larson, J.M. (2024) 'Synthetic replacements for human survey data? The perils of large language models', Political Analysis, 32(4), pp. 401–416. Available at: https://doi.org/10.1017/pan.2024.5 (Accessed: 18 August 2026).
Dai, H., Teehan, R. and Ren, M. (2024) Are LLMs prescient? A continuous evaluation using daily news as the oracle. arXiv:2411.08324. Available at: https://doi.org/10.48550/arXiv.2411.08324 (Accessed: 18 August 2026).
Fairgen (2025) Google's synthetic data study for market research. Available at: https://www.fairgen.ai/blog/google-synthetic-data-market-research-study (Accessed: 18 August 2026).
Huang, J., Li, M. and Shao, S. (2025) Distribution shift alignment helps LLMs simulate survey response distributions. arXiv:2510.21977. Available at: https://doi.org/10.48550/arXiv.2510.21977 (Accessed: 18 August 2026).
ICC/ESOMAR (2025) ICC/ESOMAR International Code on Market, Opinion and Social Research and Data Analytics. International Chamber of Commerce and ESOMAR. Available at: https://iccwbo.org/news-publications/business-solutions/iccesomar-international-code-market-opinion-social-research-data-analytics/ (Accessed: 18 August 2026).
Jia, M., Chen, Y., Sharma, D. and Diaz-Rodriguez, J. (2026) When can digital personas reliably approximate human survey findings?. arXiv:2605.10659. Available at: https://doi.org/10.48550/arXiv.2605.10659 (Accessed: 18 August 2026).
Kim, S., Wang, J., Janssen, M.A. and Anderies, J.M. (2025) How large language models systematically misrepresent American climate opinions. arXiv:2512.23889. Available at: https://doi.org/10.48550/arXiv.2512.23889 (Accessed: 18 August 2026).
Kirk, R., Mediratta, I., Nalmpantis, C., Luketina, J., Hambro, E., Grefenstette, E. and Răileanu, R. (2024) Understanding the effects of RLHF on LLM generalisation and diversity. International Conference on Learning Representations (ICLR) 2024. Available at: https://doi.org/10.48550/arXiv.2310.06452 (Accessed: 18 August 2026).
Meister, N., Guestrin, C. and Hashimoto, T. (2024) Benchmarking distributional alignment of large language models. arXiv:2411.05403. Available at: https://doi.org/10.48550/arXiv.2411.05403 (Accessed: 18 August 2026).
Moon, J., Kim, J., Lah, Y., Han, Y. and Kang, Y. (2026) Beyond averages: evaluating LLMs on human survey replication at the distributional level. arXiv:2606.09013. Available at: https://doi.org/10.48550/arXiv.2606.09013 (Accessed: 18 August 2026).
Qualtrics (2025) Qualtrics Edge. Available at: https://www.qualtrics.com/edge/ (Accessed: 18 August 2026).
Santurkar, S., Durmus, E., Ladhak, F., Lee, C., Liang, P. and Hashimoto, T. (2023) Whose opinions do language models reflect?. International Conference on Machine Learning (ICML) 2023. Available at: https://doi.org/10.48550/arXiv.2303.17548 (Accessed: 18 August 2026).
Sharma, M. et al. (2024) Towards understanding sycophancy in language models. International Conference on Learning Representations (ICLR) 2024. Available at: https://doi.org/10.48550/arXiv.2310.13548 (Accessed: 18 August 2026).
Verian Group (2026) Synthetic sample in social research: significant limitations of AI generated responses. Available at: https://www.veriangroup.com/news-and-insights/synthetic-sample-in-social-research (Accessed: 18 August 2026).
Yang, Y., Dempsey, W., Han, P., Deshmukh, Y., Richardson, S., Tom, B.D.M. and Mukherjee, B. (2024) 'Exploring the Big Data Paradox for various estimands using vaccination data from the global COVID-19 Trends and Impact Survey (CTIS)', Science Advances, 10(22). Available at: https://doi.org/10.1126/sciadv.adj0266 (Accessed: 18 August 2026).
Explore the idea
Let’s talk
Invisible forces shape your world — until you hire Latenta®
Contact