Synthetic Data for Survey Pretesting: Fake the Survey, Not the Findings

Article M2-05

You can now generate a whole survey's worth of fake respondents on demand and run your study's machinery against them before a single real person answers. That is genuinely useful for shaking down the routing, the timing and the code. It turns dangerous the moment you read what the fake answers actually say.

In brief

A survey is two things at once: a machine that routes, times and records answers, and the answers themselves. Synthetic pretesting means using fake respondents before you run the survey, to find technical problems while fixing them is still inexpensive. The honest version of the claim is narrow. Testing the machine is worth doing, but for most of it a language model is the wrong tool, because structured random data does the same job and cannot make you interpret a result that was never real. The model is useful only where the pretest needs realistic free text or a correlation pattern you cannot define in advance. One rule holds throughout. You may confirm that your analysis code runs on synthetic data. You may never trust what it says.

What the theory says

The theory

Every survey gets tested before it goes out, or it should. Piloting a survey is a long-established practice. You run a draft past a few people to check that the questions work, that the routing sends the right respondent down the right path, that the survey is not too long, and that the answers come back in a form your analysis can actually read. A practical guide to pilot testing lists the ordinary purposes: fieldwork readiness, skip logic, timing, and pre-coding the open answers (Rhoda et al., 2023). None of that is new, and none of it required a computer.

It helps to see a survey as two separate things. One is the machine: the routing rules, the piping that carries an answer from one question into the wording of the next, the timing, the export, and the code that turns raw responses into tables. The other is the contents: the answers themselves, and the relationships between them that your analysis is built to find. Pretesting has always been mostly about the machine. You are checking that the survey works before you use it with real respondents.

What changed after 2022 is that you can now fill that container with fake respondents on demand. A language model will generate a full, respondent-level dataset for your questionnaire in minutes, at almost no cost, so the whole study can be rehearsed from end to end before a single person is recruited. That capability did not exist before, and it is the reason synthetic pretesting is worth an article rather than a line in a methods textbook. The rehearsal is genuine. The question is what you are allowed to conclude from it.

The safe answer is a single boundary. Test the container, and never read the contents. A synthetic dry-run can tell you that a respondent who picks option three is correctly skipped past the next block, that the survey takes about eleven minutes, that the export lands in the right columns, and that your analysis script runs to the end without crashing. It cannot tell you anything true about what the analysis found, because the answers were invented. Confirming that the code runs is the whole legitimate use. Reading its output is the trap.

Two nearby jobs are not this one. Whether a model can catch a badly worded question, the kind a real person would misread, is comprehension testing, and that is covered in [M1-04]. Whether synthetic answers are valid enough to stand in for real ones is a separate and larger argument, covered in [M0-03], with the reason they flatten in [M2-03] and the sampling version in [M2-01]. Reading the synthetic answers to rank and narrow a field of ideas, rather than to test the instrument, is search-space reduction, covered in [M2-10]. This article is only about the machine.

For most of the machine, you do not need a language model at all, and you are safer without one. The dominant survey platform already ships a button for this. Qualtrics can generate test responses, which are random dummy answers, precisely so you can see how your dataset and your reports will look before you field, and it keeps its language-model features on a separate path (Qualtrics, no date). Random data tests the machine just as well as fake people do, and it carries no false finding that could mislead you, because nobody could mistake it for a real result.

If you need more than random noise, you can still get it without a model. A family of tools built for exactly this lets you specify the structure you want and simulate data to match it. simstudy (Goldfeld, 2026) and faux (DeBruine, 2026) let you set the correlations between variables and the factorial design directly, so you can rehearse an analysis on data that has the shape you expect. DeclareDesign goes further and lets you declare a whole research design, simulate from it, and diagnose whether the analysis recovers what you built in (Coppock, 2026). This is a well-established practice, not a new development. Peer-reviewed tools for simulating data to test an analysis predate language models by years (Regular et al., 2020; Gambarota and Altoè, 2024; Keller, 2024).

So the model's actual use is narrower than the proposal suggests. Random data covers the plumbing. Specified data covers a correlation structure you already understand. That leaves two jobs the model can do that the older tools cannot. It can produce believable free text, so you can shake down the coding pipeline for open-ended answers, and it can produce a joint structure across many variables that you could not have specified in advance. Outside those two cases, using a model adds risk you did not need.

Controversies

The current disagreement is whether synthetic respondents add anything beyond that plain baseline. Vendors say yes, and say it concretely. Running AI personas through the actual survey, rather than random data, is how you catch bugs that only appear when a realistic pattern of answers triggers your piping and skip logic (Expected Parrot, 2026; Mishkin, 2026). The claim is plausible, and it is the strongest thing the pro side has.

What it does not have is a benchmark behind it. The one study that measures the underlying capability points the other way. Chen, Zhu and Zheng (2026), in a preprint that has not yet been peer-reviewed, tested four language models against strong non-model baselines across two large social surveys and found that no model beat even the best non-model baseline at reproducing human answers. That result is about answer quality rather than survey mechanics directly, and one preprint is weak evidence, so it should be treated cautiously. But it makes the question clearer. If the fake people are not better than a simple baseline at reproducing human answers, the case for using them instead of random data rests on the free-text and unknown-structure parts, not on the whole task.

There is a second, quieter argument about whether this is a method at all. Synthetic pretesting is one of the least contested uses of synthetic data, and a good part of the reason is that the underlying craft is old. Piloting is decades old. Simulating data to test an analysis is decades old. The genuinely new ingredient is the on-demand, end-to-end rehearsal, and reasonable people can disagree about whether that ingredient is enough to make this a distinct method or just classic piloting with a modern update. The honest framing is modest: the new capability is real but narrow. Someone selling a synthetic-respondent platform will describe it more favorably, which is worth remembering when you read the pitch.

Limitations

The boundary must be strict because a synthetic dataset's relationships are untrustworthy in a specific, measured way. Bisbee et al. (2024) compared answers generated by a language model against real human survey data and found that the averages often looked plausible while the relationships underneath did not. Of the regression coefficients they examined, 48% differed significantly from the human ones and 32% flipped sign, and the model's answers were too consistent with each other, understating the variance a real sample shows. A synthetic pilot can confirm that your regression runs and fills a table. If you then read that table, roughly a third of the time the effect points the wrong way, and you have no way of knowing which third.

The false-confidence risk is not an imagined worry. It has been measured, though for statistical synthesis rather than for language models. El Emam et al. (2024) tested whether analyses run on synthetic health data reach the same conclusions as the same analyses on real data. They replicated only when at least ten synthetic datasets were generated and combined with the proper statistical rules. A single synthetic dataset, which is exactly what a quick pretest produces, was misleading. It gave wrong confidence-interval coverage and made the analysis look more powerful than it was. The danger is not that synthetic data is obviously wrong. It is that a single run looks right and reads confident while being neither.

Synthetic respondents also break the survey machinery they are meant to test, and this has two effects. A market-research firm writing about its own synthetic pilots reports personas that mangle skip logic and fake respondents that fail quality checks (Mishkin, 2026). One effect is a real limit. Fake respondents violate survey logic on complex designs, so they are a poor stand-in for careful humans. The other effect is a specific trap in the pretest itself. When a synthetic respondent trips your routing, you cannot tell whether it found a genuine bug in your skip logic or simply failed to follow the survey the way a person would. Its own logic violations can look like your routing error. This is the opposite of the false-confidence problem: a false positive to match the false negative.

The positive evidence is limited and comes from just outside the relevant time period and from a different field. Zhuang et al. (2022) showed that generating synthetic data from observed processes gives a low-cost alternative to collecting pilot data when you want to estimate statistical power. It is a clear demonstration of the comparator idea: simulate before you collect. But it is dated to the very edge of the post-2022 line, it uses generative modelling rather than a language model, and it comes from brain imaging rather than survey research. It supports the old practice of simulating to test an analysis. It does not support the new claim that a language model should do the simulating.

Open questions

Three questions the field has not answered decide how much a synthetic pretest is really worth.

Does a synthetic pilot catch the machine bugs a real pilot catches, and at what rate? As far as this research could establish, no independent study has scored whether a synthetic dry-run actually finds the broken routing, the wrong timing and the failing export that a human pilot would, or how often it misses them. Vendors assert that it does, but nobody has measured its hit-rate against a real pilot on the same instrument. The test is straightforward to run, and until someone runs it, the whole safety claim is based on reasoning and vendor demos, not on a measured result.

When a synthetic respondent trips your routing, can you tell a real bug from the model's own rule-breaking? This is the false-positive problem mentioned above, and nobody has shown how to separate the two signals: a synthetic pilot that flags a routing problem may have found a genuine fault or may simply have failed to take the survey the way a person would. Until someone can tell them apart, even the container test, the safe use, returns a result you cannot fully trust.

Where does a language model actually overtake specified or random data? The reasoned boundary is clear enough: plain data wins for well-typed numbers and known correlations, and the model is useful only for believable free text or a joint structure you cannot specify in advance. But where precisely that crossover falls has never been measured in a head-to-head trial, so a practitioner deciding whether a given pretest needs a model is making a judgement, not reading a fact.

So what

Synthetic pretesting is worth doing, it is cheaper than it has ever been, and it is safe exactly as far as the container and no further. The discipline that makes it safe is a single habit. You use synthetic respondents to prove your machine works, and you never let yourself read what they answered. If you act on a synthetic result as if it were a real finding, the tool no longer protects you and instead misleads you.

For research practice

Use the plain tool by default. For testing routing, timing, exports, and analysis code, structured random data works and cannot mislead you, so use a platform's test-response generator or a simulation package before using a model. Use the model only for the two jobs that only it can do: generating realistic open-ended text to test your coding pipeline, and generating a joint structure across variables that you could not specify by hand. Whatever you use, run the analysis to confirm it runs and fills its tables, then close the output without reading it. If you notice yourself writing that the synthetic effect was in the expected direction, stop, because that sentence is the failure. A useful check is to ask what decision the pretest serves. A pretest serves the decision to run the study or not. It never serves the decision that the study itself is meant to answer.

For companies

The sales claim is that AI lets you test your survey before you send it out, and the useful part of that claim is true. A cheap test run that catches broken piping and an overly long survey is a real benefit, especially across the many studies that today get no pretest at all. The mistake is to let the test run's answers influence the decision. When a vendor shows that its synthetic respondents match human answers closely, ask what was matched and how many datasets it took to get there, because a single synthetic run can look correct while getting a third of its relationships wrong. Treat a synthetic pretest as a way to improve the basic quality of your study, not as a preview of your results.

For political parties

A survey designed to confirm your views is not useful information. Synthetic pretesting makes it cheap to rehearse a lot of question wordings quickly and privately, which is a genuine advantage for message and issue testing, and it is also the exact setting where a fake result is most tempting to believe. Using it rigorously and responsibly means doing the same thing here. Run synthetic respondents to check that your instrument routes and times correctly across many variants, then test the surviving wordings on real people before you trust a single number about opinion. A model asked to play a voter will hand you a fluent, confident answer that a benchmark says is no better than a strong statistical baseline at being that voter (Chen, Zhu and Zheng, 2026). If you report a result based on synthetic answers, you are not learning about your coalition. You are only seeing your own assumptions.

For government and policy

Official statistics and public surveys are most affected when a system fails without being noticed, so the container test is genuinely valuable here. Testing a complex instrument's routing and its analysis pipeline before using it can find a costly mistake early. The same risks make the rule about contents mandatory. A synthetic dataset that yields an official-looking number is a risk to governance, not a preview, and the standards are changing accordingly. The ICC/ESOMAR code now separates a person from a synthetic persona and attaches disclosure and oversight duties to synthetic answers (International Chamber of Commerce and ESOMAR, 2025). The acceptable practice is to use synthetic data to test the machine, to keep a clear record of which parts of the process were tested and how, and to keep synthetic answers out of any published figure. If an analysis has to be dry-run on synthetic data, the finding is that one synthetic dataset is not enough to trust even the procedure, let alone the result (El Emam et al., 2024).

How to use this

Before you fake a survey, settle four questions. First, what are you actually testing, the machine or the answers? If it is the machine, you are on safe ground. If it is the answers, stop. Second, does this job really need a model, or will random or specified data do it more safely? Reserve the model for free text and for structure you cannot specify. Third, when the analysis runs on synthetic data, are you prepared to confirm it executes and then not read the result? If you cannot promise that, do not generate the data. Fourth, if a real decision rides on this study, have the synthetic answers been kept out of every number that reaches it? A synthetic pretest checks whether your process works, not what your results will show.

Case studies

Escalent (2026). A market-research firm published practical guidance on synthetic data and, unusually, framed synthetic pilots as experiments that can and should fail. Its account names the exact machinery failures this article is about: personas that mangle skip logic, fake respondents that fail age and quality checks, and models that overstate purchase intent (Mishkin, 2026). It is one firm's blog rather than an independent evaluation, so it is evidence of what a practitioner reports, not proof of a hit-rate. As a concrete description of what breaks when you field fake respondents against real survey logic, it is the clearest example found.

Qualtrics test responses. This is a live product proof rather than an organisation's story. The dominant survey platform ships a feature that fills a survey with random dummy answers so you can see how your data and reports will look before fielding, and it keeps its language-model generation on a separate path (Qualtrics, no date). That design choice is the deflationary argument built into the installed base. For shaking down the container, random data is the default, and the model is a different tool for a different job.

Expected Parrot / EDSL. At the open, inspectable end of the market, an open-source package invites you to run your survey on synthetic respondents before spending a real one, specifically to catch piping and skip-logic bugs on elaborate designs (Expected Parrot, 2026). It is the clearest vendor statement of the container-testing use, and it is a code package rather than a replace-your-panel pitch, which makes it a fair counterweight. It is still a vendor demonstration with no measured hit-rate, so it shows what is claimed, not what has been proven.

References

Bisbee, J., Clinton, J.D., Dorff, C., Kenkel, B. and Larson, J.M. (2024) 'Synthetic replacements for human survey data? The perils of large language models', Political Analysis, 32(4), pp. 401–416. Available at: https://doi.org/10.1017/pan.2024.5 (Accessed: 18 August 2026).

Chen, Z., Zhu, D. and Zheng, L.N. (2026) When synthetic users fail: a cross-domain benchmark of LLM-simulated human survey responses. arXiv:2607.26348, 28 July. Available at: https://arxiv.org/abs/2607.26348 (Accessed: 18 August 2026).

Coppock, A. (maintainer) (2026) DeclareDesign: declare and diagnose research designs. R package version 1.1.1, 28 April. Available at: https://cran.r-project.org/web/packages/DeclareDesign/index.html (Accessed: 18 August 2026).

DeBruine, L. (maintainer) (2026) faux: simulation for factorial designs. R package version 1.2.4, 11 March. Available at: https://cran.r-project.org/web/packages/faux/index.html (Accessed: 18 August 2026).

El Emam, K., Mosquera, L., Fang, X. and El-Hussuna, A. (2024) 'An evaluation of the replicability of analyses using synthetic health data', Scientific Reports, 14, article 6978. Available at: https://doi.org/10.1038/s41598-024-57207-7 (Accessed: 18 August 2026).

Gambarota, F. and Altoè, G. (2024) 'Understanding meta-analysis through data simulation with applications to power analysis', Advances in Methods and Practices in Psychological Science, 7(1), article 25152459231209330. Available at: https://doi.org/10.1177/25152459231209330 (Accessed: 18 August 2026).

Horton, J. / Expected Parrot (2026) Test your survey on synthetic respondents before you waste a single real one. Expected Parrot blog, 15 July. Available at: https://blog.expectedparrot.com/p/test-your-survey-on-synthetic-respondents (Accessed: 18 August 2026).

International Chamber of Commerce and ESOMAR (2025) ICC/ESOMAR international code on market, opinion and social research and data analytics. 5th edn. Paris and Amsterdam: ICC and ESOMAR. Available at: https://iccwbo.org/news-publications/business-solutions/iccesomar-international-code-market-opinion-social-research-data-analytics/ (Accessed: 18 August 2026).

Keller, B.T. (2024) mlmpower: an R package for conducting power analysis and data simulation with multilevel models. PsyArXiv/OSF preprint. Available at: https://doi.org/10.31234/osf.io/utjrg (Accessed: 18 August 2026).

Mishkin, G. (2026) Synthetic data in market research: practical guidance without the hype. Escalent blog, 2 March. Available at: https://escalent.co/blog/synthetic-data-in-market-research-practical-guidance-without-the-hype/ (Accessed: 18 August 2026).

Qualtrics (no date) Generate test responses. Qualtrics support documentation. Available at: https://www.qualtrics.com/support/survey-platform/survey-module/survey-tools/generating-test-responses/ (Accessed: 18 August 2026).

Regular, P.M., Robertson, G.J., Lewis, K.P., Babyn, J., Healey, B. and Mowbray, F. (2020) 'SimSurvey: an R package for comparing the design and analysis of surveys by simulating spatially-correlated populations', PLOS ONE, 15(5), e0232822. Available at: https://doi.org/10.1371/journal.pone.0232822 (Accessed: 18 August 2026).

Rhoda, D.A., Cutts, F.T., Agócs, M., Brustrom, J., Trimner, M.K., Clary, C.B., Clark, K., Koffi, D., Manibaruta, J.C., Sowe, A., Gunnala, R., Ogbuanu, I.U., Gacic-Dobo, M. and Danovaro-Holliday, M.C. (2023) 'A practical guide to pilot testing community-based vaccination coverage surveys', Vaccines, 11(12), article 1773. Available at: https://doi.org/10.3390/vaccines11121773 (Accessed: 18 August 2026).

simstudy: see Goldfeld, K. (maintainer) (2026) simstudy: simulation of study data. R package version 0.9.2, 9 February. Available at: https://cran.r-project.org/web/packages/simstudy/index.html (Accessed: 18 August 2026).

Zhuang, P., Chapman, B., Li, R. and Koyejo, O. (2022) Synthetic power analyses: empirical evaluation and application to cognitive neuroimaging. arXiv:2210.05835, 11 October. Available at: https://arxiv.org/abs/2210.05835 (Accessed: 18 August 2026).

Explore the idea

Let’s talk

Invisible forces shape your world — until you hire Latenta®

Contact