Survey Sample Augmentation: Borrow the Answers You're Missing, Up to a Point

Article M2-04

When a survey cell is too thin to trust, an AI model can now manufacture extra respondents to fill it out. The gains are real on the smallest cells and fade to nothing on merely small ones, and the whole thing holds only while real human answers stay in charge of the result.

In brief

A survey often has enough answers overall but too few inside a particular cell: women under twenty-five in one region, buyers of a niche product, a minor party's voters. The old fix was statistical: reweighting the people you did reach to represent the ones you did not. The new fix asks an AI model to generate the missing answers and adds them to the real ones. Done carefully, that means anchoring the synthetic answers to a held-out set of real human replies and correcting against them, never combining the two without checking. The evidence says the boost is large where the cell is tiny and disappears once the cell is only small. It also says the model reproduces group averages while reducing the variety inside them, which is why a minimum number of real human answers must control the final estimate.

What the theory says

The theory

In every survey, some subgroup has too few respondents. Say you interviewed two thousand people, which is plenty for a national figure in most countries, but only thirty-one of them were renters under thirty in the region you care about, and the estimate for that group varies so much that it is not informative. The group you most want to say something about is the one you have the least data on. This is a basic consequence of sampling, and it is the problem this whole family of methods exists to attack.

The classic answer was statistical, and the new answer uses the same logic. When a cell is thin, you can reweight the respondents you did reach so that the ones who resemble the missing group count for more, a family of techniques that runs from raking through calibration to multilevel regression and poststratification, which we cover in modelling small groups. Calibration in this older sense means forcing the survey's totals to match known population totals (Souza, Barbian and Reis, 2025). Small-area estimation uses a larger dataset to improve an estimate for a place or group where direct data is scarce (Tam and Sharmeen, 2024; Tam, 2024). And when the gap is missing values rather than missing people, imputation and oversampling fill it by using patterns from the cases you do have (Le et al., 2024; Li and Si, 2023). None of these invent respondents. They make more use of the real respondents.

The post-2022 move is to let a language model generate the missing answers directly, then splice those synthetic answers into the real ones. Early surveys of this shift catalogue several ways to do it and treat the result as an addition to the data ecosystem rather than a replacement for it (Timpone and Yang, 2024). That is a genuinely different mechanism, and people may be tempted to treat it as a replacement for asking people. It is not, and why it is not is the subject of a separate article on why synthetic respondents cannot stand in for a real sample. The question here is narrower and less discussed: once you have decided to splice rather than substitute, how is the splice actually done, and how much can you do before the estimate becomes unreliable?

There is no single method, but they share a common approach. A hybrid combines a small amount of real human data with a larger amount of model-generated data, and it is considered valid because it checks the model against real held-out human answers and corrects the model's errors. Three versions of this are visible in the field, and no source unifies them, so this is a family rather than a single algorithm. One version is statistical boosting, where a model trained on your real respondents generates extra synthetic ones for the thin cell, the approach a vendor like Fairgen sells and which is covered as a method in statistical sample boosting. A second keeps a language model in the loop but calibrates its output against real answers, as in the semantic-similarity approach we describe under simulated purchase-intent. A third treats the human budget as the scarce resource and spends it where it does the most good, sending real people the questions the model is worst at and letting the model handle the rest. What the three have in common is the base of real human answers. Each keeps a base of real human answers and treats the model as something to be corrected against that base, never as reliable on its own.

Recent method work focuses on that correction step. Ye and Yoganarasimhan (2026) formalise the allocation problem, assigning limited human labels to the items where the model is least reliable and reporting error reductions of around eleven percent in their tests. Tan and Zrnic (2026) show how to get statistically valid conclusions out of a synthetic-plus-real mix, on the condition that the synthetic task behaves like the historical real-data tasks you calibrated on. Both are preprints, which is important. The one peer-reviewed anchor for the general idea of humans and language models working as collaborators rather than substitutes is Arora, Chakraborty and Nishimura (2025) in the Journal of Marketing, which describes the hybrid as a division of labour rather than a transfer of work.

Controversies

Does adding synthetic answers to a small group give real information, or does it just repeat the model's assumptions as if they were data? This is the main disagreement, and Chen, Zhu and Zheng (2026) argue most strongly for the negative view. Testing language models as stand-in survey respondents across domains, they found that no model beat the strongest simple baseline on individual answers, that estimates of cross-cultural values ran eleven to twenty-two points worse than a plain demographic predictor, and that the models over-determined identity, letting a single attribute like a political leaning appear to explain a large share of a response where the real-world share is tiny. Their most uncomfortable finding is that larger, more capable models stereotyped more, not less. If that holds, then a boosted group can contain patterns that come from the model's assumptions about those people, not from the people themselves. The synthetic answers would look consistent because they are a simplified, exaggerated version.

The next disagreement contradicts a natural assumption, that more information should make the splice safer and that richer, newer models should splice better. Sometimes they do the opposite. Morris, Leff and Enns (2025), working at the survey firm Verasight, tested language-model imputation of survey answers and found that adding more prompt detail, adding personas, voting history and step-by-step reasoning, could lower accuracy rather than raise it, with the error on immigration attitudes reaching more than eleven points and several subgroup errors above ten. Newer flagship models did not reliably fix this. This is a problem for anyone selling a more complex synthetic method: a richer prompt is not automatically a safer one.

The camps also disagree about where the limit actually is. The industry describes the limit in terms of real sample size: a vendor states a minimum real base and a boost factor, and independent testing by Dig Insights (2026) reports the same shape, with large gains on the smallest cells and nothing left once the real cell is above roughly a hundred and fifty to two hundred people. The academic framing is different. Huang, Wu and Wang (2025) see it as a balance rather than a sudden cutoff: too few synthetic answers and you still have too much random error, too many and the estimate becomes more precise but around a guess that is not well supported, so there is a best mix somewhere in between. Both camps agree there is a limit. They do not agree on the factor that determines it, real sample size or synthetic share, and nobody has published a curve that settles it.

Limitations

The most important limitation is about the evidence, not the method. There is no independent, non-vendor, peer-reviewed replication of vendor-style thin-cell boosting. As far as this research could establish, the only independent-looking validation of that approach comes from a consultancy study and a practitioner pilot, both of which are hosted on the vendor's own website (Dig Insights, 2026; Nguyen, 2025). That does not make them wrong, and the consultancy states it has no commercial tie to the vendor, but a result you can only read on the seller's domain is weaker evidence than the same result in a journal or on a neutral site. The independent academic evidence that does exist sits almost entirely on the sceptical side, testing where synthetic answers fail (Chen, Zhu and Zheng, 2026; Morris, Leff and Enns, 2025).

The evidence is therefore thinner than the market's confidence suggests. The method-mechanics and operating-envelope claims rest on preprints that have not yet cleared peer review (Huang, Wu and Wang, 2025; Ye and Yoganarasimhan, 2026; Tan and Zrnic, 2026; Chen, Zhu and Zheng, 2026). The standards and vendor material is exactly that, evidence of what is claimed and required, never of whether a given tool works (ICC/ESOMAR, 2025; ESOMAR, no date; Fairgen, no date). The one genuinely peer-reviewed hybrid-method paper is Arora, Chakraborty and Nishimura (2025), and the rest of the peer-reviewed pool is the older calibration and small-area work that predates the language-model version. The field is barely more than a year old in its current form.

Open questions

Three questions the field has not settled determine how much you can use a boosted cell.

What is the limit, and how is it measured? No one has published the curve that relates the share of synthetic answers in a cell to the accuracy of the resulting estimate. Theory gives a shape, an interior optimum where too little synthetic gives you too much noise and too much makes the estimate too dependent on a guess with poor coverage (Huang, Wu and Wang, 2025), and a separate result tells you which real answers to use first (Ye and Yoganarasimhan, 2026). But the vendors describe the limit in real sample size while the academics describe it in synthetic share, and no one has determined which measure is correct (Dig Insights, 2026). Until something does, any dose-response curve, including the one on this page, shows how the method works, not a measured outcome, and a practitioner has no published number to use as a limit.

How much synthetic is too much, officially? No standards body has published a numeric ceiling. ESOMAR describes the limit in words, not numbers, through a Minimum Viable Data threshold and a taxonomy of augmentation methods, not as a maximum synthetic share you must not cross (ESOMAR, no date), and the ICC/ESOMAR code requires only that synthetic respondents be disclosed and kept distinct from real people, which is a rule about labels, not a limit on amount (ICC/ESOMAR, 2025). A practitioner who needs a threshold to defend in an audit does not yet have an official one to cite.

Is this one method or three? Statistical boosting, language-model calibration and optimal allocation use the same basic approach, measure against real data and leave the difficult tasks to humans, but no study has tested all three together and compared them. Whether they turn out to be one technique or three matters to a buyer choosing between vendors who each sell a different one, because there is no direct evidence to distinguish them.

So what

Methodological rigor and ethical rigor are the same thing. When you reweight real respondents, you can only ever redistribute answers people actually gave. When you splice in synthetic answers, you can create answers nobody gave, and the same tool that fills a small group honestly can be adjusted to show any result you want. The final dataset looks the same in both cases. So a useful way to think about this method is to imagine being responsible for the final number, defending an estimate for a group where most answers were made up by your organisation, not collected from people. A boosted cell is only useful if it was calibrated against real held-out answers and kept within the range where the boost genuinely improves accuracy.

For research practice

Treat augmentation as a way to improve a cell with few responses, not as a way to reduce the sample size. The approach supported by the evidence is to ensure every reported cell includes some real human responses, use the model to boost only where the cell is genuinely tiny, and always calibrate the synthetic output against a held-out set of real answers so you can measure the model's error before trusting it. Use your real respondents where the model is least reliable instead of distributing them equally (Ye and Yoganarasimhan, 2026). And follow the limits: the benefits occur only for the smallest cells and disappear when a cell is just small, so improving a cell that was already adequate adds risk without benefit (Dig Insights, 2026). The failure to watch for is the subtle one. A boosted estimate that looks clean and plausible is the least informative result, because a model asked to fill a cell will almost always produce something clean and plausible, whether or not it matches the real people.

For companies

You will be told that synthetic augmentation lets you cut sample cost while keeping your subgroup readouts. The truth is that it lets you get a usable read on the tiny segments you could never afford to sample properly, which is genuinely useful, as long as you do not also stop sampling the segments you can afford. Brand and concept tests live on subgroup cuts, and the smallest cuts are exactly where a synthetic number that looks good is most appealing and least verified. So when a vendor reports an improvement, ask two things. Ask what the synthetic answers were calibrated against, because a boost based on your own real respondents is more valuable than one based on a generic model. And ask where you can read the validation, because the evidence that looks independent currently appears only on the vendors' own websites (Dig Insights, 2026; Nguyen, 2025). One such pilot, on a single platform, found average confidence-interval gains of around a quarter across dozens of brand-lift cuts, with the biggest gains on segments under fifty people (Nguyen, 2025). A number you cannot trace to a neutral source is a claim, not a result.

For political parties

The manipulation risk here is not hypothetical. A party that wants a particular subgroup to look a particular way, young men breaking toward it, a region that was undecided starting to support it, can now boost a thin cell until it shows exactly that result, and the boosted subgroup result looks the same as a real one. That is the survey-research version of a push poll, a tool designed to create a conclusion instead of discover one. The difference between research and persuasion is whether the boost helps you make a decision or just confirms what you already believed. If your campaign acts on a boosted subgroup, it is spending real resources on people whose views were mostly created by a model that, on the independent evidence, tends to over-determine identity and caricature the very groups you are trying to understand (Chen, Zhu and Zheng, 2026). Used well, augmentation is genuinely valuable here, because it lets you narrow a large field of message and issue variants quickly and privately before you spend on real fieldwork. Where researchers have checked language-model estimates of vote choice against real election results, the model matched the overall pattern but was unreliable for smaller subgroups (von der Heyde, Haensch and Wenz, 2025), which is the same warning at national scale. The careful approach is to use the model to narrow down options and then test the remaining ones on real voters from the groups that matter, instead of letting the simulation represent a group whose real votes will later disagree with it. A survey designed to agree with you is useless as information.

For government and policy

Official statistics are the most important, because a synthetic answer that reaches a published figure becomes a number used to make real decisions, and the groups a model has most trouble reproducing are often the groups a public survey is meant to cover. The older calibration and small-area methods behind this work have a long, defensible track record in official statistics (Tam and Sharmeen, 2024; Souza, Barbian and Reis, 2025), and the new synthetic version should be required to meet at least that standard, not a weaker one. Two rules matter most. The first is disclosure: the profession's code now requires that synthetic respondents be labelled and kept distinct from real people, and an agency should record which cells were boosted, by how much, and against what real data they were calibrated (ICC/ESOMAR, 2025). The second is honesty about limitations. Where a boosted estimate cannot be validated against real answers for the group it describes, that limitation should be stated in the release, not in a footnote. A statistical agency that publishes where the method failed, rather than only where it succeeded, is the one whose numbers remain trustworthy.

How to use this

Before you accept a boosted estimate, ask four questions. Is there a real human floor in this cell, or is the number almost entirely synthetic? What were the synthetic answers calibrated against, and how wrong was the model when it was checked against real held-out replies? Is this cell small enough that boosting plausibly helps, or was it already large enough that the boost adds risk for nothing? And does filling this cell serve the decision in front of you, or the conclusion someone was hoping to reach? If you cannot answer those, treat the cell as thin and say so. The method's real advance is that it finally makes a defensible estimate possible for groups too small to sample well, which is worth a great deal. It has not made the real human answer optional.

Case studies

Verasight (2025). In the clearest independent test of the sceptical case, researchers at the survey firm Verasight used language models to impute survey answers and build synthetic samples, then checked the results against real responses. Adding more information to the model, personas, voting history and step-by-step reasoning, did not reliably help and sometimes hurt, with the error on immigration attitudes reaching more than eleven points and several subgroup errors above ten (Morris, Leff and Enns, 2025). The value of the study is that it measures the failure where it matters most, at the subgroup level, which is exactly where thin cells and boosting live.

Dig Insights and Fairgen (2026). For the positive pole, a consultancy study reported that statistical boosting roughly tripled the effective reliability of the smallest cells, with the largest gains on real bases of ten to forty people and essentially no benefit once a cell reached a hundred and fifty to two hundred, including a small regional example rising from around fifteen to roughly eight in its margin (Dig Insights, 2026). The consultancy states it has no commercial tie to the vendor, but the study is published on the vendor's own website, and no peer-reviewed, fully independent replication exists yet.

References

Arora, N., Chakraborty, I. and Nishimura, Y. (2025) 'AI–Human Hybrids for Marketing Research: Leveraging Large Language Models (LLMs) as Collaborators', Journal of Marketing, 89(2), pp. 43–70. Available at: https://doi.org/10.1177/00222429241276529 (Accessed: 18 August 2026).

Chen, Z., Zhu, D. and Zheng, L.N. (2026) When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses. arXiv:2607.26348 [preprint]. Available at: https://doi.org/10.48550/arXiv.2607.26348 (Accessed: 18 August 2026).

Dig Insights (2026) Independent validation of Fairgen Boost. Published 14 January 2026. Available at: https://www.fairgen.ai/blog/synthetic-data-validation-independent-study (Accessed: 18 August 2026).

ESOMAR (no date) 5 Topics of Discussion to Help Buyers of Augmented Synthetic Data. Available at: https://esomar.org/uploads/attachments/cmf43exde4helqr3vcld9txhv-esomar-5-topics-of-discussion-to-help-buyers-of-augmented-synthetic-data.pdf (Accessed: 18 August 2026).

Fairgen (no date) Boost [product page]. Available at: https://www.fairgen.ai/platform/boost (Accessed: 18 August 2026).

Huang, C., Wu, Y. and Wang, K. (2025) How Many Human Survey Respondents is a Large Language Model Worth? An Uncertainty Quantification Perspective. arXiv:2502.17773 [preprint]. Available at: https://doi.org/10.48550/arXiv.2502.17773 (Accessed: 18 August 2026).

ICC/ESOMAR (2025) ICC/ESOMAR International Code on Market, Opinion and Social Research and Data Analytics. 5th edn. Available at: https://iccwbo.org/news-publications/business-solutions/iccesomar-international-code-market-opinion-social-research-data-analytics/ (Accessed: 18 August 2026).

Le, H.H. et al. (2024) Oversampling and imputation for imbalanced missing data. Research Square [preprint]. Available at: https://doi.org/10.21203/rs.3.rs-4562263/v1 (Accessed: 18 August 2026).

Li, K. and Si, Y. (2023) 'Embedded multilevel regression and poststratification: Model-based inference with incomplete auxiliary information', Statistics in Medicine, 43(2), pp. 256–278. Available at: https://doi.org/10.1002/sim.9956 (Accessed: 18 August 2026).

Morris, G.E., Leff, B. and Enns, P.K. (2025) The Limits of Synthetic Samples in Survey Research. Verasight, 26 September. Available at: https://www.verasight.io/reports/synthetic-sampling-2 (Accessed: 18 August 2026).

Nguyen, M. (2025) Google / ESOMAR pilot on augmented synthetic data [practitioner pilot]. Available at: https://www.fairgen.ai/blog/google-synthetic-data-market-research-study (Accessed: 18 August 2026).

Souza, J.S. de, Barbian, M.H. and Reis, R.C.P. dos (2025) 'Comparison of calibration methods in the analysis of 2013 Brazilian National Health Survey data', Revista Brasileira de Epidemiologia, 28. Available at: https://doi.org/10.1590/1980-549720250005 (Accessed: 18 August 2026).

Tam, S.-M. (2024) 'Small data estimation for binary variables with big data: A comparison of calibrated nearest neighbour and hierarchical Bayes methods of estimation', Statistical Journal of the IAOS, 40(3), pp. 581–589. Available at: https://doi.org/10.3233/sji-240007 (Accessed: 18 August 2026).

Tam, S.-M. and Sharmeen, S. (2024) 'A calibrated data-driven approach for small area estimation using big data', Australian & New Zealand Journal of Statistics, 66(2), pp. 125–145. Available at: https://doi.org/10.1111/anzs.12414 (Accessed: 18 August 2026).

Tan, L. and Zrnic, T. (2026) Valid Inference with Synthetic Data via Task Exchangeability. arXiv:2606.13629 [preprint]. Available at: https://doi.org/10.48550/arXiv.2606.13629 (Accessed: 18 August 2026).

Timpone, R. and Yang, Y. (2024) Artificial Data, Real Insights: Evaluating Opportunities and Risks of Expanding the Data Ecosystem with Synthetic Data. arXiv:2408.15260 [preprint], presented at IC2S2. Available at: https://doi.org/10.48550/arXiv.2408.15260 (Accessed: 18 August 2026).

von der Heyde, L., Haensch, A.-C. and Wenz, A. (2025) 'Vox Populi, Vox AI? Using Large Language Models to Estimate German Vote Choice', Social Science Computer Review, 44(3), pp. 549–571. Available at: https://doi.org/10.1177/08944393251337014 (Accessed: 18 August 2026).

Ye, Z. and Yoganarasimhan, H. (2026) Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys. arXiv:2604.17267 [preprint]. Available at: https://doi.org/10.48550/arXiv.2604.17267 (Accessed: 18 August 2026).

Explore the idea

Let’s talk

Invisible forces shape your world — until you hire Latenta®

Contact