The Say–Do Gap in AI Research: A Machine Learns What People Say, Not What They Do

Article M0-07

The oldest problem in survey research just got a new kind of respondent: one that has never once acted on a preference.

In brief

For decades the first rule of reading a survey has been that what people say they will do is a weak guide to what they actually do. Now the industry has a respondent that is made entirely of what people have said. Large language models are trained on text, and text is the record of stated preference, not behaviour. So when a synthetic respondent gives you an answer, it is answering on the say side of a gap the field has spent fifty years learning to distrust. The uncomfortable part is how the tools get validated: their answers are usually checked against other survey answers, then sold as a prediction of behaviour. The fix has not changed. Ground the claim in what people did, whether the respondent is a person or a model, before you bet on it.

What the theory says

The theory

In people, this problem is old and well measured. Ask them what they intend, and their answers only loosely track what they go on to do. In the classic review of the field, stated intentions explain a little over a quarter of the variation in later behaviour (Sheeran, 2002). Push people's intentions up hard in an experiment, and their behaviour still barely moves (Webb and Sheeran, 2006). The pattern is old enough to have a name in economics. A preference you read off a choice someone actually made is a revealed preference. A preference someone only states is a different thing. Samuelson drew that line back in 1938.

The gap is not just real. It has been measured. Some studies ask people what they would pay, then check against what they actually pay. The hypothetical answer overstates real willingness to pay by a median factor of about 1.35, and sometimes by much more (Murphy et al., 2005). The same inflation turns up inside the format market research leans on most, the choice experiment, where people pick between designed options (Hensher, 2010). None of this is a reason to stop asking. It is the reason serious research treats a stated answer as a starting point and then corrects it. Some methods warn people about the tendency to overstate before they answer. Others apply correction factors built from past gaps between what people said and what they did. Every one of those fixes works the same way. It ties the stated answer back to observed behaviour, because there is no other way to calibrate a claim about behaviour. We cover the human side of this in our companion piece on the say-do gap in people. The profession knew all of it long before any of this was automated. Saying is not doing, and the only reliable cure is to check the saying against the doing.

Now put a language model in the respondent's chair. A large language model is trained to predict text, and the text it learns from is overwhelmingly the record of what people have said: reviews, forums, survey archives, social posts, marketing copy. It is a compression of stated things. That is the insight behind using models as synthetic respondents. The idea is usually traced to the silicon sampling work of Argyle et al. (2023), who showed a model could reproduce the average opinions of a real group closely enough to pass for a sample of it, and we cover it properly in silicon sampling. The consequence for the say-do gap is simple, and easy to miss. A model of people built from text is a model of what people say. It has no independent access to what they did.

The machine version is not just different from the human one. It is quietly worse. A real person is unreliable about their intentions, but they also act, and some of those actions can be seen. That gives you something to check the stated answer against. A model gives you nothing of the kind. It was built from the record of stated things, so it holds only what people wrote, never what they did. It does not just inherit the human say-do gap. It is made of the say side of the gap, and the do side was never in it.

The marketing evidence bears this out. Reviewing the studies that compare synthetic and human samples, Sarstedt et al. (2024) find that "results vary considerably across different domains", and conclude that the honest place for silicon samples is early on, in pretesting and pilot work, not the main study a decision rests on. The tool is real and useful. It helps in the early part of research, when you are still forming questions. That is exactly the part that does not yet claim to predict behaviour.

Controversies

The sharpest evidence comes from asking models directly for a preference and checking the answer against how people actually choose. When Goli and Singh (2024) tested whether a model could recover human preferences over time, the models were systematically more impatient than people, with discount rates well outside the human range, and reasoning prompts narrowed the gap without closing it. Their conclusion is blunt for anyone planning to elicit preferences straight from a model: doing so can produce misleading results. This is the say-do gap arriving on cue. The model states its preference fluently, in complete sentences, with reasons attached. And it can be wrong in a way you would only catch by comparing it to behaviour. Fluency is the trap. A person who is unsure will hedge, go quiet, or tick don't know, and that hesitation is information. A model rarely does that. It returns a confident, well-formed answer to almost any question, so the usual signs that a stated preference is soft go missing, exactly when you need them most.

Willingness to pay tells the same story, with a twist that matters commercially. Prompted to act as consumers, models produce demand curves that slope the right way and price sensitivities that look plausible next to human studies, at least in familiar categories (Brand, Israeli and Ngwe, 2024). The trouble sits in the structure underneath the average. The synthetic answers vary too little from person to person, miss the differences between groups, and grow unreliable in unfamiliar categories. That last case is the exact place a company most needs a read it cannot get any other way. A number that looks right on average can sit on top of a distribution that is wrong. It is the same failure covered in detail in synthetic validity, where synthetic answers match the average and lose the spread. What the say-do angle adds is this: even where the machine's stated numbers look sensible, they are still stated numbers, and no amount of surface realism turns a statement into an observed act.

There is also good reason to doubt that a model's survey answers reflect any stable attitude at all. Bisbee et al. (2024) found that a model prompted to stand in for public opinion matched real averages but was, in their words, "not reliable for statistical inference": the answers varied less than real ones, shifted with small changes in prompt wording, and drifted over a few months. Other work shows a model's survey answers moving with question order and format, in ways that look like quirks of the machine rather than signals from a population (Dominguez-Olmedo, Hardt and Mendler-Dünner, 2023). If the answer changes when you rephrase the question, you are measuring the instrument, not the respondent.

This is the part a buyer has to watch, because it is where the say-do gap slips past everyone. Vendors of synthetic respondents almost all advertise that they predict behaviour, and almost all back the claim with a number that measures something else. The validation is usually the similarity between the model's answers and a panel of human survey answers. One widely marketed tool reports 85 to 92 percent parity with human responses, which is parity with what people said, not with what they did (Synthetic Users, no date). That compares one stated preference to another. It is say measured against say, then sold as a forecast of doing. A few firms position themselves against exactly this. One markets its product on the line that people are unreliable narrators of their own behaviour, and that a simulation of them will act rather than merely answer (Aaru, no date). The ambition is the right one. But as far as this review could establish, the public evidence for it is a set of after-the-fact case studies, not a controlled test that predicted behaviour in advance and then checked it against the market.

Limitations

In fairness, synthetic respondents can be made to work, and the conditions that make them work prove the rule rather than break it. The strongest recent result comes from a method that stops asking the model for a rating and instead reads meaning out of its free text, then calibrates the whole system against thousands of real human responses (Maier et al., 2025). Anchored that way, across dozens of personal-care studies, it reaches about 90 percent of the reliability you would get from the same person answering twice, and it reproduces product rankings. It is a genuine achievement and a genuine commercial tool. It is also a clean illustration of the rule, not an exception to it. It works because it is tied to a real distribution of human answers, and even then it is checked against stated purchase intent rather than against purchases. It tracks the average better than the edges.

The one study that reaches all the way to behaviour reaches only part of the way. Validating model-generated brand perceptions against real car trade-in choices, Li et al. (2024) found better than three-quarters agreement on which brands sit near each other. That is real behaviour entering the test, and the model doing respectably against it. It is also similarity at the aggregate, on how a market clusters brands, not a prediction of which car a given person drives off the lot. Models also reproduce famous behavioural results from economics, fairness and status-quo effects and the like, but qualitatively, as directions rather than magnitudes you could bank (Horton, 2023). Reproducing the shape of a known finding is not the same as forecasting a number nobody has measured yet.

Open questions

The largest gap is one of absence. As far as this research could establish, no published study directly compares a model's stated preferences against real individual purchase or transaction data at scale. Validation runs almost entirely stated-against-stated, model output against human survey answers. The lone study that brings real behaviour in does so at the aggregate.

Two more absences follow from it. Nobody has yet shown, in a controlled test, that grounding a model in behaviour actually fixes the gap on fresh cases. Vendor decks and professional guidance assert the cure, and it is the right instinct. But the experiment that would prove it has not surfaced: calibrate on behaviour, then predict behaviour you have not seen. And the central claim here, that a model trained on text can only give you the say side, rests on the pattern of results rather than on a single measured number. It is the simplest explanation for why machines match survey answers and miss behaviour, which is not the same as proof. The field is young enough that saying so is a finding, not a hedge.

So what

For research practice

Treat a synthetic answer exactly as you already treat a human survey answer: as a weak signal that has to be grounded in behaviour before anyone bets on it. This is not a new discipline you have to learn for a new tool. It is the discipline you already owe every stated preference, applied to a respondent that happens to be a model. The one rule that changes everything is what you validate against. If a synthetic result has only been checked against other survey answers, you have measured whether the model can imitate a questionnaire, which it can. You have learned nothing yet about behaviour. Keep synthetic respondents where the evidence puts them, in the upstream work of forming questions and pressure-testing designs, and keep a real behavioural holdout as the referee for anything a decision rests on.

For companies

The line to hold with a vendor is simple and it is the one they least want asked. Against what real behaviour was this validated? If the answer is a similarity score against a human survey, that is say-versus-say, and it does not support a claim about what your customers will do. Buy the tool for what it is good at. It is fast and cheap for the early, exploratory part of a project, and it beats a blank page for pilots and pretesting. Do not let it make the go or no-go call on a launch on its own.

One temptation here turns research into theatre. A synthetic study is unusually easy to steer. The same model gives different answers when you change the prompt, the persona, or the framing. So a team that already knows the answer it wants can go looking for the recipe that returns it, and the tooling will oblige. That is not research. It is an expensive way to hear your own assumptions read back to you. The rigorous move and the honest move are the same move: fix the method and the behavioural benchmark before you run anything, and let the number fall where it falls. A study built to agree with you is worthless as intelligence, and the market proves you wrong anyway when the product ships. You pay for research to be told something you did not want to hear, while it is still cheap to hear it.

For political parties

A synthetic electorate is the most seductive version of this, and the most dangerous. It is cheap, it never tires, and it will answer any question about any district in seconds. It is also made of what voters have said, and that is the weakest possible basis for predicting the one thing that matters: whether they turn out. When one team tested models against a real European election, the models predicted an average turnout of 83 percent against an actual 49 percent, a miss large enough to wreck any plan built on it (reported in Ochoa, 2025). Stated enthusiasm is not turnout. A model of stated enthusiasm is enthusiasm squared. Use synthetic tools to draft and stress-test message ideas if you like. Then measure real behaviour, canvass returns, early-vote files, field experiments, before you spend a dollar or a door-knock.

For government and policy

Government has leaned on stated-preference methods for a long time, especially to put a value on things that have no market price, from clean air to a public service, by asking people what they would pay. That method already carried a known hypothetical bias before any of it was automated. A synthetic respondent makes it worse. Now you are valuing a public good with a model of what people say they value, which sits two steps away from anything anyone did. The defensible use is auditability. Professional standards now stress that AI-assisted methods have to be transparent and open to audit, and that a synthetic result should be checked against real data rather than taken on trust (ESOMAR, 2025). For a public decision, that is not a nicety. A number that will justify spending or regulation has to trace back to behaviour someone can check. It cannot rest on a model's fluent guess about behaviour it never saw.

How to use this

One question does most of the work, whoever is selling you the number. What behaviour was this validated against, and can I see it? If the honest answer is that the synthetic answers were compared to other survey answers, you are looking at a stated preference wearing the costume of a prediction. Treat it as a hypothesis to test against behaviour, not as the test. Then apply the oldest check in the field to the newest respondent in it: find the place where you can watch what people actually did, and see whether the machine's confident answer survives the contact.

Case studies

GPT and willingness to pay (Brand, Israeli and Ngwe, 2024). Prompted to act as consumers, models produced demand curves and price sensitivities that looked reasonable in familiar categories, then lost the differences between groups and grew unreliable in unfamiliar ones. A clean demonstration that a plausible stated answer can sit on top of a structure that would not survive a real purchase.

Semantic Similarity Rating, PyMC Labs and Colgate-Palmolive (Maier et al., 2025). The method reads meaning out of the model's free text and calibrates against roughly 9,300 real human responses across 57 studies. Done that way, it reached about 90 percent of human test-retest reliability and recovered product rankings. The counter-case that proves the rule: it works because it is anchored to real human answers, and it is still validated against stated intent rather than purchases.

New Coke (1985). The human baseline, from before any of this was automated. Around 200,000 blind taste tests stated a clear preference for the sweeter formula, the product launched, and the market rejected it within weeks. The tests measured taste in isolation and missed the brand attachment that drove the actual behaviour (Encyclopaedia Britannica, no date). Saying is not doing, and it never was.

References

Aaru (no date) Aaru: simulate populations. Available at: https://aaru.com/ (Accessed: 13 August 2026).

Argyle, L.P., Busby, E.C., Fulda, N., Gubler, J.R., Rytting, C. and Wingate, D. (2023) 'Out of one, many: using language models to simulate human samples', Political Analysis, 31(3), pp. 337–351. Available at: https://doi.org/10.1017/pan.2023.2 (Accessed: 13 August 2026).

Bisbee, J., Clinton, J.D., Dorff, C., Kenkel, B. and Larson, J.M. (2024) 'Synthetic replacements for human survey data? The perils of large language models', Political Analysis, 32(4), pp. 401–416. Available at: https://doi.org/10.1017/pan.2024.5 (Accessed: 13 August 2026).

Brand, J., Israeli, A. and Ngwe, D. (2024) 'Using GPT for market research', in Proceedings of the 25th ACM Conference on Economics and Computation (EC '24). New York: ACM, p. 613. Available at: https://doi.org/10.1145/3670865.3673479 (Accessed: 13 August 2026).

Dominguez-Olmedo, R., Hardt, M. and Mendler-Dünner, C. (2023) Questioning the survey responses of large language models. arXiv:2306.07951. Available at: https://arxiv.org/abs/2306.07951 (Accessed: 13 August 2026).

Encyclopaedia Britannica (no date) New Coke. Available at: https://www.britannica.com/topic/New-Coke (Accessed: 13 August 2026).

ESOMAR (2025) ICC/ESOMAR International Code on Market, Opinion and Social Research and Data Analytics. Available at: https://esomar.org/icc-esomar-code (Accessed: 13 August 2026).

Goli, A. and Singh, A. (2024) 'Frontiers: can large language models capture human preferences?', Marketing Science, 43(4), pp. 709–722. Available at: https://doi.org/10.1287/mksc.2023.0306 (Accessed: 13 August 2026).

Hensher, D.A. (2010) 'Hypothetical bias, choice experiments and willingness to pay', Transportation Research Part B: Methodological, 44(6), pp. 735–752. Available at: https://doi.org/10.1016/j.trb.2009.12.012 (Accessed: 13 August 2026).

Horton, J.J. (2023) Large language models as simulated economic agents: what can we learn from homo silicus? NBER Working Paper 31122. Cambridge, MA: National Bureau of Economic Research. Available at: https://doi.org/10.3386/w31122 (Accessed: 13 August 2026).

Li, P., Castelo, N., Katona, Z. and Sárváry, M. (2024) 'Frontiers: determining the validity of large language models for automated perceptual analysis', Marketing Science, 43(2), pp. 254–266. Available at: https://doi.org/10.1287/mksc.2023.0454 (Accessed: 13 August 2026).

Maier, B.F., Aslak, U., Fiaschi, L., Rismal, N., Fletcher, K., Luhmann, C.C., Dow, R., Pappas, K. and Wiecki, T.V. (2025) LLMs reproduce human purchase intent via semantic similarity elicitation of Likert ratings. arXiv:2510.08338. Available at: https://arxiv.org/abs/2510.08338 (Accessed: 13 August 2026).

Murphy, J.J., Allen, P.G., Stevens, T.H. and Weatherhead, D. (2005) 'A meta-analysis of hypothetical bias in stated preference valuation', Environmental and Resource Economics, 30(3), pp. 313–325. Available at: https://doi.org/10.1007/s10640-004-3332-z (Accessed: 13 August 2026).

Ochoa, C. (2025) 'Synthetic respondents and the future of survey research', Quirk's Media, 21 October. Available at: https://www.quirks.com/articles/synthetic-respondents-and-the-future-of-survey-research (Accessed: 13 August 2026).

Samuelson, P.A. (1938) 'A note on the pure theory of consumer's behaviour', Economica, 5(17), pp. 61–71. Available at: https://doi.org/10.2307/2548836 (Accessed: 13 August 2026).

Sarstedt, M., Adler, S., Rau, L. and Schmitt, B.H. (2024) 'Using large language models to generate silicon samples in consumer and marketing research: challenges, opportunities, and guidelines', Psychology & Marketing, 41(6), pp. 1254–1270. Available at: https://doi.org/10.1002/mar.21982 (Accessed: 13 August 2026).

Sheeran, P. (2002) 'Intention–behavior relations: a conceptual and empirical review', European Review of Social Psychology, 12(1), pp. 1–36. Available at: https://doi.org/10.1080/14792772143000003 (Accessed: 13 August 2026).

Synthetic Users (no date) Synthetic Users: user research, without the users. Available at: https://www.syntheticusers.com/ (Accessed: 13 August 2026).

Webb, T.L. and Sheeran, P. (2006) 'Does changing behavioral intentions engender behavior change? A meta-analysis of the experimental evidence', Psychological Bulletin, 132(2), pp. 249–268. Available at: https://doi.org/10.1037/0033-2909.132.2.249 (Accessed: 13 August 2026).

Explore the idea

Let’s talk

Invisible forces shape your world — until you hire Latenta®

Contact