Synthetic Respondents for Hypothesis Generation: A Hypothesis Engine, Not a Truth Machine
Article M2-10
Point a synthetic panel at two hundred ideas and it will hand back the best twelve in an afternoon, for almost nothing. Whether those are the twelve a real test would have chosen is the one thing nobody has independently shown, and the whole case for the method rests on it.
In brief
Synthetic respondents are language models prompted to stand in for a sample of real people, and their one defensible job is search-space reduction: screening a large field of ideas, concepts or messages down to a shortlist that real people then test. Give a panel two hundred concepts and it returns a ranked dozen in minutes, at almost no cost. The whole case for that funnel rests on one property, rank preservation, meaning the shortlist the synthetic screen passes is the same shortlist a real test would have passed. That has been measured once, by a vendor and its client, and never confirmed by anyone without a stake. Everything the method reliably does, reproduce broad patterns and the direction of a result, supports screening; everything it reliably fails at, the spread of opinion, the missing tails, the numbers that change with the prompt and the model, disqualifies it as the measurement a decision rests on. This is not dry-running a survey's plumbing, which is [M2-05] and never reads the answers; here you read them, but only to rank and narrow, never to decide.
What the theory says
The theory
A synthetic respondent is a language model asked to answer as a person, or a sample of people, would. You give it a profile, a questionnaire or a set of concepts, and it produces the answers it predicts that kind of person would give. The word synthetic has come to mean two unrelated things in this corpus, which is why we also wrote the article [M2-11] to keep them apart. For this article, it means a simulated respondent standing in for a human one. It does not mean the older idea of synthetic data, where you statistically generate records that preserve the shape of a real dataset without exposing anyone in it.
The technique is genuinely new. Argyle et al. (2023), a peer-reviewed paper that named the idea, showed that conditioning a model on a real person's background and attitudes produced answers that tracked the matching group of Americans. They called this algorithmic fidelity, meaning the model reproduces the patterns of a known human group closely enough to be worth something. That paper is the origin of the category, and the original claim it made about using simulated people is the subject of [M2-01]. It is best read as the origin rather than as proof that the method is ready to replace fieldwork. It showed the thing was possible. It did not show where it is safe.
The defensible use has a shape and a name: search-space reduction. Imagine a funnel. Two hundred candidate ideas go in at the wide end, a fast, cheap synthetic screen ranks them and cuts the field to a shortlist of a dozen, real people test the survivors, and a decision comes out at the narrow end. The synthetic panel does one job in that funnel: it reads the ideas and puts them in order, so the expensive human test runs only on the candidates that are worth testing. This is not testing the survey instrument itself, the routing, timing and analysis code, which is the subject of [M2-05]. The scaffolding is old, innovation management has staged filters under the name stage-gate for decades (Tavares-Quinhoes and Velez-Lapão, 2023); what the post-2022 evidence adds is a clear idea of where in that funnel a simulated respondent is useful, and one hard question about whether it truly does.
The strongest statement of it comes from Sarstedt et al. (2024), a peer-reviewed piece in Psychology and Marketing. They found particular promise for simulated samples in the upstream parts of a project, qualitative pretesting and pilot studies, and cautioned against leaning on them for the main study that a decision depends on. They also found that how well the method works varies a great deal from one domain to another, which matters later. In a market-research setting specifically, Brand, Israeli and Ngwe (2024), a peer-reviewed conference paper, found that a model reproduced well-known economic patterns: demand that slopes downward as price rises, and sensible willingness-to-pay, which is the highest price a person will accept before refusing to buy. So the model has a real grasp of the broad shape of how people behave.
That grasp is why it works well as a filter but not as a source of final evidence. The most useful habit this area offers costs nothing: before you run a simulation, state clearly which slot you are in. Are you generating candidates to test, or are you trying to settle a question? Stating that first is required. Almost every misuse in the market is treating a filter result as a decision.
Controversies
The key question, and the one the evidence has not answered, is whether a synthetic screen preserves rank order. Rank order is just the sequence a test would put your candidates in, best to worst. A filter is only useful if the candidates it passes are the ones a real test would also have passed. If the synthetic screen fails the eventual winner or passes an eventual loser, it has not saved you a step. It has added a wrong step.
The strongest peer-reviewed support is Ashokkumar et al. (2026), a paper in Nature, which found that language models could anticipate the direction of results across a large set of social-science experiments. Direction means which way an effect goes, not how big it is and not where a given option lands in a ranking. Predicting direction is a real and useful capability, and it is the strongest reason to trust a synthetic filter at all. It is also a weaker claim than rank preservation. The closest thing to direct evidence appeared at the end of 2025. Maier et al. (2025), a preprint written by a synthetic-data consultancy with its client Colgate-Palmolive, tested language models on 57 real personal-care concept surveys and found that they reproduced the ranking of concepts by average purchase intent at about 90 percent of the reliability you get when the same people answer twice. That is rank order on real concepts, which is the important claim. It is also a single preprint, not peer-reviewed and not repeated by anyone without a stake in the result, so it changes the evidence from indirect to a single study by an interested party.
There is a more serious concern behind this. Gui and Toubia (2023), a working paper and so a step below peer review, argue that an underspecified prompt leaves the simulated experiment confounded. In plain terms, if you do not pin down the details, the model can produce a directional answer that reflects an artifact of how you asked rather than any real human cause. Two independent 2026 preprints now put numbers on that worry. Chen, Zhu and Zheng (2026) found that models treat demographics as far more telling than they really are, inflating the gaps between groups several times over and pointing a decision at the wrong group in about half the cases they tested. Lukauskas and Šarkauskaitė (2026) found that answers modelled on synthetic data lost their predictive power against real people. Both are a step below peer review, and both land in the same place. That makes the friendly label risky. Calling an output directional, use with care sounds cautious, but it can also launder a confounded result into something a team feels licensed to act on. The label reassures people more than the evidence justifies.
Limitations
Downstream, as the measurement itself, the method fails in a way that is now well documented. Bisbee et al. (2024), peer-reviewed in Political Analysis, found that synthetic survey answers can produce a plausible-looking average while the distribution is wrong. The range of opinions narrows, and the extreme views disappear, and the estimates are unstable: change the prompt wording, the model version, or even the date you run it, and the numbers change. An average you cannot reproduce next week is not a valid measurement.
The reason this happens, and the one partial repair for it, are taught elsewhere in this series of articles. The mechanism is flattening, the model's tendency to move answers toward the middle of a distribution and away from the extremes, which [M2-03] and [M2-04] cover, with the underlying evidence in [M0-03]. The repair is calibration, correcting synthetic output against a small real anchor. What matters here is only the consequence for the funnel. Flattening is acceptable in a filter, because a filter needs the broad shape and not the tails. It makes the measurement invalid, because the extreme values are often the answer.
Two more limits sit around the method. The domain-dependence Sarstedt et al. (2024) found means a filter that works for one category can quietly fail in the next, so success with snack foods does not mean success with pharmaceuticals. And the commercial evidence is weaker than the marketing claims. Vendor accuracy claims cluster between about 85 and 92 percent (Synthetic Users, 2026), but under different names, similarity here, parity there, correlation somewhere else, with no common metric and, as far as this research could establish, no disclosed holdout of real answers for a fair test. Subconscious.ai reports the most specific public benchmark, its synthetic answers matching a human baseline on a well-known immigration conjoint study at roughly 0.83 against a human-to-human ceiling of about 0.96. Evidenza states 88 percent accuracy across more than a hundred validations, supported by testimonials, not a published method. One notable detail: the same client engagement has been claimed by two different vendors, which shows how validation is done in this market. A vendor document shows what is claimed. It does not show that the claim is true.
That was the full situation until 2026, when the first independent check was published. STRAT7, an analytics group that sells human research rather than a synthetic product, ran a nationally representative sample of 3,000 real people against surveys from two synthetic-data firms. The synthetic answers came in two to three points off the human headline, put prices in an illogical order 68 percent of the time, and tracked change between two waves just 47 percent of the time, no better than chance. An independent benchmark now exists, and it supports the evidence of failure, not the vendors' claims.
There is a rule now, but it is not useful. The revised ICC/ESOMAR International Code, approved by the profession's two main bodies in 2025, makes disclosure mandatory (Research World, 2026): if synthetic data or AI shaped a piece of research, the audience must be told, along with how much human oversight there was. That rule covers honesty about using the tool, not whether the numbers it produced are accurate. On that second question, as far as a targeted search could establish, no professional body has published a certification, accuracy benchmark or audit protocol for synthetic respondents. ESOMAR's own ESOMAR 24 is a questionnaire for vetting human sample suppliers, not a synthetic-data certification, and the Insights Association's July 2026 webinar series is education. A duty to say you used the tool and a standard for whether it works are different, and only the first exists. That missing accuracy standard is the risky part of the funnel, because the place a method is easiest to misuse is the place where no one has agreed on what good performance is.
One last limit affects the real-people step that the funnel depends on. Westwood (2025), peer-reviewed in PNAS, warns that language models now contaminate online survey panels with machine-generated answers, which makes the human test that is supposed to catch the simulation's mistakes more expensive and less reliable. The check at the bottom of the funnel is getting more expensive while the simulation at the top is getting cheaper.
Open questions
Three questions the field has not answered decide whether the funnel is a demonstrated method or a well-reasoned bet.
Does the ranking hold in a genuinely new category, or only where the model has already seen the field? The screen looks strongest on familiar concepts a model has read a great deal about, and whether it preserves rank order for a novel product, a thin category, or a non-Western audience, exactly where a study provides its value, has not been tested. A screen that ranks well only on easy ideas is not very useful, so the transfer to the hard cases is the gap a buyer most needs addressed.
Does the hybrid actually beat a human-only process on cost at equal decision quality? The main economic argument is that a synthetic filter plus a smaller human test is cheaper than fielding humans alone for the same quality of decision, and a targeted search found no measured study of that trade-off. In the STRAT7 test a blended sample roughly halved the price-ordering errors of a pure-synthetic one, but that is a blend beating pure simulation on one error rate, not a hybrid beating human-only on cost.
When a synthetic screen is wrong, does it mostly reject good concepts or accept bad ones? Those two errors require opposite safeguards, and no study settles which dominates. A buyer who does not know cannot tell whether to guard against a lost winner or a false pass, which is the difference between trusting the filter too much and using it too little.
So what
The theory has one rule and one bet. The rule: a synthetic panel is for search-space reduction, cutting a large field of ideas to a shortlist real people then test, and never for the measurement a decision rests on. The method depends on the bet that the shortlist the screen gives you is the shortlist a real test would have chosen, and that has been shown once by interested parties and confirmed by no one. The more costly it is to be confidently wrong, the earlier in the funnel the simulation has to stay, and the more that unconfirmed bet should matter to you.
For research practice
Treat synthetic respondents as a first filter and never as the finding. The sequence that fits the evidence is to run a wide field of concepts, wordings or segments through a simulated panel early, use it to cut the field to a shortlist, and then spend your real fieldwork on testing the survivors. That is where the method is strong: reproducing broad patterns and the direction of a result, which is enough to drop the obvious losers cheaply. What it has not been independently shown to do is preserve rank order, so when you test the survivors on real people, check that they really are the ones the screen ranked highest; that check is the only way to learn whether the funnel held for your category. The discipline is to write down, before you run it, that this is a screen. A screen produces candidates to test, not answers to report. The mistake to avoid is when a directional filter result is later presented as a firm finding, with the word directional removed. And remember that the method is domain-dependent, so a filter you trust in one category has to prove itself again in the next.
For companies
The pitch you will hear is that synthetic respondents let you skip expensive fieldwork. The honest version is that they let you afford more of the early exploration you skip today, which is a real gain, while the replacement story is where the money gets lost. A concept test run only on simulated buyers returns clean-looking numbers with the disagreement flattened out, and the product still has to meet real customers who were never asked. When a vendor reports that its synthetic answers match humans at 88 or 90 percent, ask three things: matched against what real data, held back in what holdout, and measured how. Until 2026 no independent benchmark of any commercial tool existed; the first, from STRAT7, cut against the sales pitch rather than for it. So any single accuracy number is a claim and not a result. Aaru's wealth-research work is an example of what can go wrong, covered in [M2-06]. Use the simulation to widen what you explore. Keep real customers as the gate on anything you are going to spend against.
For political parties
This is where a simulated respondent is most easily misused. A synthetic panel will produce whatever answer the prompt leans toward, and unlike a real focus group, no participant in it can push back and no independent benchmark can catch it. That makes it tempting to use it to create a false impression of public support: run the model until the message tests well, then present the output as evidence that the public agrees. The check every question has to pass is whether the exercise serves the decision you actually face or the conclusion the campaign already preferred. A poll designed to agree with you is useless as information, because you have paid for something that only reflects your own views. Using it rigorously and responsibly is the same thing here. Use the simulation to narrow a large field of wordings and issues fast and in private, which it does well, and then test the survivors on real voters, especially the parts of a coalition a fluent model is least able to speak for, the hard-to-reach audiences [M2-07] takes up.
For government and policy
The same ethical rule applies more strictly here, because the output can become an official number that guides a decision, and the people most harmed by a bad tool are often the people the work is meant to help. A simulated respondent can widen coverage cheaply at the exploratory stage, helping a team scope questions or pressure-test a design before fielding it. It must never be the number of record. Using the method for final measurements causes the most harm, hiding the differences across groups that policy should consider. The acceptable approach is a synthetic first step that expands what is studied, paired with real data collection for anything that drives a decision, and a clear written record of which findings came from a simulation and which from people. The decision test applies when public money is used: does this analysis help the decision, or just support a conclusion a sponsor wanted? Publishing the method, including where the simulation fell short, is the honest move, and it also prevents the tool from secretly making a preferred answer look like a statistic.
How to use this
Four habits carry the whole argument.
Name the slot before you run it. Write down whether this simulation is generating candidates or settling a question. If it is settling a question, stop, because that is the job it cannot do.
Treat every synthetic output as a hypothesis, not a result. A directional finding is a reason to test something, never a reason to skip the test.
Keep a real-people gate on anything that drives a decision or a spend. The simulation widens what you look at. People still decide what is true.
Ask any vendor the same three questions: matched against what real data, held back in what holdout, measured how. If the answer is a testimonial, you have learned what is claimed and nothing about whether it holds.
Case studies
PyMC Labs and Colgate-Palmolive (2025). The most concrete named deployment is a preprint, not a vendor landing page. The consultancy PyMC Labs, working with its client Colgate-Palmolive, ran language models against 57 of the brand's real personal-care concept surveys covering 9,300 people, and reported that the models reproduced the ranking of concepts by average purchase intent at roughly 90 percent of the score you get when the same people retake the survey. It is the strongest evidence in this article that a synthetic screen can preserve rank order on real concepts, and it was written by the tool's seller and buyer, so it is a serious and disclosed demonstration by interested parties, not independent proof.
Subconscious.ai and the immigration conjoint (2026). The most specific public benchmark in the market comes from Subconscious.ai, which reports that its synthetic respondents reproduced the results of a well-known immigration conjoint study at a correlation of about 0.83, against a human-to-human ceiling of roughly 0.96, and an average near 0.73 across dozens of studies. Read as a Tier C vendor figure, it is useful in two ways. It is the clearest illustration of how good a synthetic filter can be, and, because it is vendor-run with no independent holdout disclosed, it is also the clearest illustration of what passes for validation in this market. The number is real evidence of what is claimed. It is not evidence of the claim.
Fairgen's digital-twin boost (2026). Fairgen offers the closest commercial version of the two-stage hybrid the funnel implies: augmenting a small real sample with synthetic respondents to stretch it further, with output the company describes as directional, not definitive. Its reported 98 percent sanity pass is internal quality assurance, not external accuracy, and the underlying figures are self-reported. It is included here as the market's nearest analogue to the cost-versus-quality hybrid the evidence has not yet measured, not as a verified case that the hybrid works.
Aaru and EY wealth research (2026). Handled in full in [M2-06], and noted here only for the one detail that belongs to this article. The same EY engagement has been cited by more than one synthetic-data vendor, which is a compact picture of how validation-by-testimonial works. When the proof of a method is a client story that competing vendors both claim, the proof is doing marketing, not measurement.
References
Argyle, L.P., Busby, E.C., Fulda, N., Gubler, J.R., Rytting, C. and Wingate, D. (2023) 'Out of one, many: using language models to simulate human samples', Political Analysis, 31(3), pp. 337–351. Available at: https://doi.org/10.1017/pan.2023.2 (Accessed: 18 August 2026).
Ashokkumar, A., Hewitt, L., Ghezae, I. and Willer, R. (2026) 'Large language models can predict the results of social science experiments', Nature, 656(8126), pp. 115–122. Available at: https://doi.org/10.1038/s41586-026-10742-x (Accessed: 18 August 2026).
Bisbee, J., Clinton, J.D., Dorff, C., Kenkel, B. and Larson, J.M. (2024) 'Synthetic replacements for human survey data? The perils of large language models', Political Analysis, 32(4), pp. 401–416. Available at: https://doi.org/10.1017/pan.2024.5 (Accessed: 18 August 2026).
Brand, J., Israeli, A. and Ngwe, D. (2024) 'Using GPT for market research', in Proceedings of the 25th ACM Conference on Economics and Computation (EC '24). New York: ACM, p. 613. Available at: https://doi.org/10.1145/3670865.3673479 (Accessed: 18 August 2026).
Chen, Z., Zhu, D. and Zheng, L.N. (2026) 'When synthetic users fail: a cross-domain benchmark of LLM-simulated human survey responses', arXiv. Available at: https://doi.org/10.48550/arXiv.2607.26348 (Accessed: 19 August 2026).
ESOMAR (2026) ESOMAR. Available at: https://www.esomar.org/ (Accessed: 18 August 2026).
Evidenza (2026) Evidenza. Available at: https://evidenza.ai/ (Accessed: 18 August 2026).
Fairgen (2026) Fairgen. Available at: https://www.fairgen.ai/ (Accessed: 18 August 2026).
Gui, G. and Toubia, O. (2023) 'The challenge of using LLMs to simulate human behavior: a causal inference perspective', SSRN Electronic Journal. Available at: https://doi.org/10.2139/ssrn.4650172 (Accessed: 18 August 2026).
Insights Association (2026) Insights Association. Available at: https://www.insightsassociation.org/ (Accessed: 18 August 2026).
Lukauskas, M. and Šarkauskaitė, V. (2026) 'Plausible but not valid: a psychometric audit of LLMs as synthetic survey respondents', arXiv. Available at: https://doi.org/10.48550/arXiv.2608.14606 (Accessed: 19 August 2026).
Maier, B.F., Aslak, U., Fiaschi, L., Rismal, N., Fletcher, K., Luhmann, C.C., Dow, R., Pappas, K. and Wiecki, T.V. (2025) 'LLMs reproduce human purchase intent via semantic similarity elicitation of Likert ratings', arXiv. Available at: https://doi.org/10.48550/arXiv.2510.08338 (Accessed: 19 August 2026).
Research World (2026) 'Why the ICC/ESOMAR Code will matter more than ever in 2026', Research World, 8 January. Available at: https://researchworld.com/articles/why-the-icc-esomar-code-will-matter-more-than-ever-in-2026 (Accessed: 19 August 2026).
Sarstedt, M., Adler, S.J., Rau, L. and Schmitt, B. (2024) 'Using large language models to generate silicon samples in consumer and marketing research: challenges, opportunities, and guidelines', Psychology & Marketing, 41(6), pp. 1254–1270. Available at: https://doi.org/10.1002/mar.21982 (Accessed: 18 August 2026).
STRAT7 (2026) Synthetic data: is this as good as it gets? Reported in 'Synthetic data survey responses significantly different to humans', Research Live, July 2026. Available at: https://www.research-live.com/article/news/synthetic-data-survey-responses-significantly-different-to-humans/id/5151157 (Accessed: 19 August 2026).
Subconscious.ai (2026) Subconscious.ai. Available at: https://www.subconscious.ai/ (Accessed: 18 August 2026).
Synthetic Users (2026) Synthetic Users. Available at: https://www.syntheticusers.com/ (Accessed: 18 August 2026).
Tavares-Quinhoes, T.A. and Velez-Lapão, L. (2023) 'Strengthening the innovation management: insights from the stage-gates model', Journal of Technology Management & Innovation, 18(2), pp. 91–105. Available at: https://doi.org/10.4067/s0718-27242023000200091 (Accessed: 18 August 2026).
Westwood, S.J. (2025) 'The potential existential threat of large language models to online survey research', Proceedings of the National Academy of Sciences, 122(47). Available at: https://doi.org/10.1073/pnas.2518075122 (Accessed: 18 August 2026).
Explore the idea
Let’s talk
Invisible forces shape your world — until you hire Latenta®
Contact