Synthetic Voters: Can a model stand in for the electorate?

Article P1-10

Three of seven battleground calls were right, and the model's creator shrugged. What can an AI that plays a voter actually measure?

In brief

A language model prompted to answer as a voter can land near a national poll average, but it fails the checks that make a poll usable. Its answers vary far too little between people, move when the question is reworded, and lean left on party and ideology items. It also produces almost no don't knows, so it can flip which side leads. The profession therefore treats synthetic respondents as a planning tool, not a measurement: useful for testing how a question lands, not for saying what the public thinks.

How to use this

Before quoting a synthetic poll, ask what it was for: testing wording or generating hypotheses is what it can do. Check the spread, not just the average, and ask whether answers vary between respondents at all. Ask what share answered don't know, and treat a near-zero share as a warning. Re-run the same prompt with different wording and after a delay, and see whether the answer moves. Look at how the model stands on party and ideology items before trusting its read on politics. Where a figure will decide something, use a survey of real people.

What the story is about

Aaru builds AI voter avatars. It takes census and demographic data, creates a persona for each voter, feeds the personas a diet of news, then polls them. On 4 November 2024, with the avatars fed news up to midnight, the company published its probabilities for the seven US battleground states. It favoured Trump in Arizona at 73.3%, North Carolina at 62.1% and Georgia at 61.8%, and Harris in Michigan at 63.3%, Nevada at 53.4%, Pennsylvania at 52.4% and Wisconsin at 50.9% (Chua, 2024). Three of the seven calls named the eventual winner; Trump won all seven (Federal Election Commission, 2025). Two days later a co-founder defended the method: "A coin flip is a coin flip". The avatars, he said, were "significantly faster and cheaper than traditional polling, and still more accurate" (Mendoza, 2024).

Consensus

A synthetic respondent is a language model prompted with a persona: you are a homeowner in Ohio, and here is what you think. Bisbee et al. (2024) did this at scale, prompting ChatGPT to rate 11 sociopolitical groups on a scale from cold to warm (0 to 100) and checking the answers against the 2016 to 2020 American National Election Study. The average scores lined up closely with the survey averages. Everything else failed. The synthetic answers varied less than the real ones, the relationships between answers differed significantly from the survey's, minor changes in wording moved the distribution, and the same prompt produced significantly different results over a three-month period. The authors conclude that "sampling by ChatGPT is not reliable for statistical inference". Demographics alone do not repair any of this (Von der Heyde et al., 2024).

A survey reports more than an average. It also reports a spread: some people strongly agree, some mildly, some disagree. The synthetic version gets the average roughly right and the spread badly wrong, because its answers clump tightly around a single point. Boelaert et al. (2025) call this machine bias and describe it as "a strong bias and a low variance on each topic". The bias also varies randomly between topics. That is what makes the problem structural rather than a correctable quirk. An adjustment learned on one topic does not carry to the next, because the direction of the error has moved. Real electorates disagree with themselves. This one does not.

Controversies

On politics, the error has a direction. Four independent studies across three continents find a leftward lean on party and ideology items. Santurkar et al. (2023) built OpinionsQA from polls covering 60 US demographic groups and found misalignment as large as the Democrat to Republican divide on climate change, even when they steered the model toward a group. Two German studies found the same ordering, best for Green and Left supporters and worst for the right-wing AfD (Von der Heyde, Haensch and Wenz, 2025; Ma et al., 2025). The American Association for Public Opinion Research (2026) reports liberal-viewpoint bias as one of the prevailing findings across models and countries. In the 2024 US cycle, Parikh, Cen and Podimata (2026) found every model overpredicted favourability for Harris by 10 to 40% against five high-quality polls, with smaller and less consistent errors for Trump.

Some studies report striking successes. Argyle et al. (2023) prompted GPT-3 with thousands of sociodemographic backstories and produced silicon samples whose answers tracked real subgroups closely, a property they named "algorithmic fidelity". Sanders, Ulinich and Schneier (2023) found the model anticipated both the level and the spread of US opinion on policy issues cheaply. Yu et al. (2024) validated an election-prediction pipeline against the 2016 and 2020 American National Election Study. González-Bustamante, Verelst and Cisternas (2025) tested 128 prompt, model and question combinations against a Chilean probability survey and reached accuracy above 0.90 on trust items. The catch is timing. Almost every positive result compares the model with a survey or an election that happened before its training cut-off, so the model may be recalling the answer rather than reasoning to it. The Chilean study is the hard case: it is not in English, and it worked. It also measured trust, not vote choice.

Limitations

The errors concentrate where the groups are small. Morris (2025), in a Verasight white paper testing five OpenAI models against 1,500 American adults surveyed in June 2025, found average absolute error on the whole-sample figures of 4 to 23 percentage points; the best model averaged 4. Across demographic subgroups the average gap was 8 points, with Black respondents 15 points off and other-race respondents 20 points off. Then there is the answer that never comes. On Trump approval, the synthetic sample's don't know share rounded to 0% against 3% in the real data. On a zoning and housing-supply question it could not have seen, the real sample split 28% for limiting local zoning rules, 43% against and 29% don't know. The synthetic sample returned 59% for, 41% against and 0% don't know. It swapped the leading position and deleted a third of the electorate's uncertainty.

The profession's guidance is explicit. The American Association for Public Opinion Research (2026), the body that sets survey standards, states that a synthetic response is "a model-based approximation of what a person might say, not a direct observation of human expression", and that such responses "pose particularly serious validity and disclosure risks if applied beyond clearly labeled pretesting, pilot work, or exploratory diagnostics". Park et al. (2026) built agents from two-hour interviews, structured surveys or both, and those agents matched between 82% and 86% of how consistent each person's own answers were when the same questions were put to them twice, two weeks apart. Demographics-only agents, the basis of a synthetic poll, matched 74%; the higher figures came only when the individuals' own answers had been collected by asking them. The literature contains no pre-registered, prospective synthetic poll of a named future election that has been fielded before the vote and scored against the result.

A synthetic respondent is a language model asked to act like a voter, and it can land close to a national poll average. That is the whole of what it does well. It fails the checks that make a poll usable: too little variation between people, results that move when the wording moves, a leftward tilt on party and ideology items, and no real don't know. So the profession treats it as a planning tool, not a measurement. A correct average can sit on top of a wrong distribution, and the distribution is what you needed.

So what

Synthetic polls are useful for exploring language, testing how a question lands, and generating hypotheses cheaply. They cannot be trusted as measurements of public opinion. For anyone who acts on a number, that difference is the whole point. A poll you cannot check is not a cheaper poll. It is a guess with a decimal point.

For political parties

Accuracy is not uniform. Qu and Wang (2024) tested ChatGPT against World Values Survey data and found it performed best in Western, English-speaking, developed nations, notably the United States, with further gaps by gender, ethnicity, age, education and social class, and thematic biases in political and environmental topics. For a party, that means a synthetic poll is not a measure of its actual electorate. It can still be useful for pretesting how a message reads or mapping the arguments an opponent might use. Treating the model's answers as your voters' answers is the mistake.

For government

Models are better at describing an opinion distribution than at reproducing one. Meister, Guestrin and Hashimoto (2025) benchmarked distributional alignment and found exactly that split. A government can therefore use these tools to explore the range of views on a question, and to see which positions exist and how they relate. It cannot use them to claim a mandate. Setting policy on synthetic numbers, or announcing that the public supports something because a model said so, mistakes a description of opinion for a measure of it.

References

American Association for Public Opinion Research (2026) Responsible AI integration in survey research. Alexandria, VA: AAPOR, May. Available at: https://aapor.org/wp-content/uploads/2026/05/Responsible-AI-Integration-In-Survey-Research.pdf (Accessed: 9 September 2026).

Argyle, L.P. et al. (2023) 'Out of one, many: using language models to simulate human samples', Political Analysis, 31(3), pp. 337–351. Available at: https://doi.org/10.1017/pan.2023.2 (Accessed: 9 September 2026).

Bisbee, J. et al. (2024) 'Synthetic replacements for human survey data? The perils of large language models', Political Analysis, 32(4), pp. 401–416. Available at: https://doi.org/10.1017/pan.2024.5 (Accessed: 9 September 2026).

Boelaert, J. et al. (2025) 'Machine bias: how do generative language models answer opinion polls?', Sociological Methods & Research, 54(3), pp. 1156–1196. Available at: https://doi.org/10.1177/00491241251330582 (Accessed: 9 September 2026).

Chua, G. (2024) 'An AI polling startup makes its predictions for the 2024 US election', Semafor, 4 November. Available at: https://www.semafor.com/article/11/04/2024/an-ai-polling-startup-polls-bots-predicts-harris-will-win (Accessed: 9 September 2026).

Federal Election Commission (2025) Official 2024 presidential general election results. Washington, DC: FEC. Available at: https://www.fec.gov/introduction-campaign-finance/election-results-and-voting-information/ (Accessed: 9 September 2026).

González-Bustamante, B., Verelst, N. and Cisternas, C. (2025) Emulating public opinion: a proof-of-concept of AI-generated synthetic survey responses for the Chilean case. Empiria Lab Method Series. Available at: https://doi.org/10.5281/zenodo.17077752 (Accessed: 9 September 2026).

Ma, B. et al. (2025) 'Algorithmic fidelity of large language models in generating synthetic German public opinions: a case study', in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Vienna: Association for Computational Linguistics, pp. 1785–1809. Available at: https://doi.org/10.18653/v1/2025.acl-long.90 (Accessed: 9 September 2026).

Meister, N., Guestrin, C. and Hashimoto, T. (2025) 'Benchmarking distributional alignment of large language models', in Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Albuquerque, NM: Association for Computational Linguistics, pp. 24–49. Available at: https://doi.org/10.18653/v1/2025.naacl-long.2 (Accessed: 9 September 2026).

Mendoza, D. (2024) 'AI polling company defends wrong predictions on the US election', Semafor, 6 November. Available at: https://www.semafor.com/article/11/06/2024/ai-startup-aaru-defends-using-artificial-intelligence-for-polling (Accessed: 9 September 2026).

Morris, E.G. (2025) Your polls on ChatGPT. Verasight White Paper Series, 18 August (page updated 1 September 2026). Available at: https://www.verasight.io/reports/synthetic-sampling (Accessed: 9 September 2026).

Parikh, R., Cen, S.H. and Podimata, C. (2026) Do LLMs track public opinion? A multi-model study of favorability predictions in the 2024 U.S. presidential election. arXiv:2602.06302. Available at: https://doi.org/10.48550/arXiv.2602.06302 (Accessed: 9 September 2026).

Park, J.S. et al. (2026) LLM agents grounded in self-reports enable general-purpose simulation of individuals. arXiv:2411.10109v3, revised 28 June 2026. Available at: https://doi.org/10.48550/arXiv.2411.10109 (Accessed: 9 September 2026).

Qu, Y. and Wang, J. (2024) 'Performance and biases of large language models in public opinion simulation', Humanities and Social Sciences Communications, 11(1)., 1095. Available at: https://doi.org/10.1057/s41599-024-03609-x (Accessed: 9 September 2026).

Sanders, N.E., Ulinich, A. and Schneier, B. (2023) 'Demonstrations of the potential of AI-based political issue polling', Harvard Data Science Review, 5(4). Available at: https://doi.org/10.1162/99608f92.1d3cf75d (Accessed: 9 September 2026).

Santurkar, S. et al. (2023) Whose opinions do language models reflect? arXiv:2303.17548. Available at: https://doi.org/10.48550/arXiv.2303.17548 (Accessed: 9 September 2026).

Von der Heyde, L. et al. (2024) United in diversity? Contextual biases in LLM-based predictions of the 2024 European Parliament elections. arXiv:2409.09045, posted 29 August 2024 (version 2, 17 April 2025, is the version read). Available at: https://doi.org/10.48550/arXiv.2409.09045 (Accessed: 9 September 2026).

Von der Heyde, L., Haensch, A.-C. and Wenz, A. (2025) 'Vox populi, vox AI? Using large language models to estimate German vote choice', Social Science Computer Review, 44(3), pp. 549–571. Available at: https://doi.org/10.1177/08944393251337014 (Accessed: 9 September 2026).

Yu, C. et al. (2024) Towards more accurate US presidential election via multi-step reasoning with large language models. arXiv:2411.03321, posted 21 October 2024 (version 3, 4 April 2025). Available at: https://doi.org/10.48550/arXiv.2411.03321 (Accessed: 9 September 2026).

Explore the idea

Let’s talk

Invisible forces shape your world — until you hire Latenta®

Contact