Modelling Who Will Vote: The electorate that didn't show up as planned

Article P1-05

The same 2,424 people gave pollsters leads from two points Democratic to seven points Republican. The rule for counting them, not their answers, decided which.

In brief

The rule a poll uses to decide who will vote can move its headline by more than the gap it is measuring. It is a choice, not a fact about any respondent, and different rules applied to the same answers produce opposite elections. A screen trained on who voted last time has nothing to work with for a first-time voter, so it can undercount a group that turns out in numbers, as happened in 2024 and in the UK's 2024 general election. Modelled screens did somewhat better on average among firms that answered, but that is an association across 39 self-selected firms, not proof, and no single method guaranteed accuracy.

How to use this

Before trusting a poll's headline, ask what rule it used to decide who will vote. Ask whether that rule was trained on past voters, and if so, how it scores someone with no voting history. Check whether the screen's likely movement is larger than the gap being reported; if it is, treat the lead as provisional. When a firm says it used a voter file, note that it may also weight on partisanship, so the effect of the screen alone is hard to isolate. For campaign planning, remember that screening out low-propensity voters removes people you may be trying to reach. If a poll does not disclose its likely-voter method, discount the precision of its number.

What the story is about

One survey produced sixteen different answers. Pew Research Center took a single set of 2,424 people, matched them to a national voter file so it could check what they actually did, and ran the same answers through sixteen different rules for deciding who would really turn out. The results ranged from a 2-point Democratic lead to a 7-point Republican advantage. The benchmark they were graded against was a 3-point Republican lead among verified voters. Nothing about the respondents changed between those results. Only the rule did. The study is a methods report, not a peer-reviewed paper, and it is a decade old. It is still the best teaching case the subject has, because the alternatives were published alongside the answer (Keeter and Igielnik, 2016).

A likely-voter rule is a sorting device, and a person chooses it. Each respondent is scored, the ones who clear the bar have their answers counted, and everyone else is set aside, however they answered. The classic version of that rule was built at Gallup by Paul Perry across the 1950s and 1960s, and Perry (1973) published the comparison of what likely voters and likely nonvoters preferred. The index asks seven questions covering four things: how much thought someone has given the election, whether they have voted before, how closely they follow public affairs, and how likely they say they are to vote this time. None of those questions is a fact about November. Together they are a bet about who will show up.

Every one of those sixteen rules beat the naive alternatives. Counting everyone who says they are registered to vote, or everyone who says they intend to vote, was worse than any screen, because those groups include far too many people who never cast a ballot (Keeter and Igielnik, 2016). So the decision to screen is easy. Choosing the screen is not, and different screens pull the headline in different directions. Two pollsters can survey the same people, publish honestly, and describe opposite elections. Pollsters face the same kind of fork when they adjust a sample to match the population's demographics and partisanship, which is the subject of how polls are weighted.

Researchers have tried to build a better rule. Rentsch, Schaffner and Gross (2019) combined stated intention with demographic predictors of turnout and of overreporting, the habit of saying you voted when you did not. Their peer-reviewed conclusion is that the bias and error created by likely-voter models can be reduced to a negligible amount, using variables pollsters already collect. The caution comes from a different corner. Ansolabehere et al. (2024) fitted the leading academic explanations of turnout to past elections and used them to predict the next one, and reported that such saturated models overfit the data and lead to less accurate predictions than parsimonious models. Their subject is the turnout level in a place, not which individuals vote, so this is no direct refutation. It is a warning against assuming that a bigger model is a better one.

In 2024 the modelled route did better on average, at least among the firms that answered. Among those firms, polls built on multivariate turnout models or registration-based screens had somewhat lower average errors than polls relying on simple self-reported likelihood-of-voting questions (American Association for Public Opinion Research, 2025). The report states the difficulty itself. Firms that used voter files also tended to weight on partisanship and to use detailed likely-voter models, which makes the effect of any single decision hard to isolate. With 39 self-selected firms, that is an association, not proof that modelling caused the improvement. And no single methodological choice guaranteed more accurate results.

Turnout in 2024 moved in both directions at once. "Republican turnout surged in rural and exurban counties, while Democratic turnout fell in some urban centers", and the shift bit hardest in Georgia, Pennsylvania and Nevada. Most polls assumed the new electorate would look like the last one, and only a few built local geography or regional turnout projections into their weights. That is the pattern the task force describes. The mechanism underneath it is simpler. A turnout model learns from who voted last time, so it has nothing to work with for a person who has never voted. In 2024 that mattered: 2020 nonvoters who turned out were undercounted, many of them Republican-leaning, and "their turnout likelihood was often underestimated by models trained or selected on habitual election participants" (American Association for Public Opinion Research, 2025).

The rule a poll uses to decide which respondents will actually vote can move its headline number by several points, more than the gap it is measuring. That rule is a choice, and no choice is neutral. A screen trained on who voted last time has nothing to say about a person who is voting for the first time, and in 2024 there were enough of those people to matter. Both sets of arithmetic can be correct. The picture of the electorate can still be wrong.

So what

The likely-voter decision reaches past the mechanics of polling. It can change the story a poll tells, and it matters most in the elections where the story is least settled, when turnout patterns shift and new voters arrive in numbers. Parties use polls to decide where to spend, governments use them to judge public opinion, and everyone else reads the headline. A published number that looks like a measurement carries a judgement inside it about whose answers counted.

For political parties

For a party, a poll number is a judgement call as much as a measurement. Screening out people who look unlikely to vote removes, by design, some of the low-propensity and first-time voters a campaign may be trying to reach. The record the models lean on is itself unstable. Kim and Fraga (2022), in a preprint, found that a voter file gives a different picture of who voted depending on when the snapshot was taken, and that "low-propensity voters are particularly impacted". Where no such file exists to buy, the turnout probability has to be built from demographics and self-report instead, because the commercial files are made from official, publicly available government records of who is registered and who has voted (Pew Research Center, 2018). The risk is in the method.

For government

The worry that a turnout screen skews the result is not new. The 2016 committee found evidence of a turnout shift favouring Trump, but judged the evidence that misspecified likely-voter models caused the error to be mixed (Kennedy et al., 2018). The 2020 task force tested the screen and found it sorted supporters of both candidates about equally well; its heading on that question reads "Error Due to Method of Voting and Likely Voters: Unlikely to Be a Factor". Its next heading kept likely-voter modelling on the suspect list (American Association for Public Opinion Research, 2021). The 2024 report is sharper; what remains unknown is the size of the error, whether the modelling can be separated from the firms that use it, and whether the lesson holds for a midterm.

Case studies

In the UK, the same failure turned up without any voter file to lean on. More in Common polled the 2024 general election and then published what its turnout model got wrong: young Green voters, scored as unlikely to turn out because that is how people like them had behaved before. Its own account says the model uses age because younger people are less likely to vote, and adds that "in this election we underestimated the Green Party's performance in part because many young Green voters were modelled as unlikely to vote, in line with previous elections" (More in Common, 2024). Two countries, opposite political directions, one failure mode.

References

American Association for Public Opinion Research (2021) 2020 Pre-Election Polling: an evaluation of the 2020 general election polls. Washington, DC: American Association for Public Opinion Research. Available at: https://aapor.org/wp-content/uploads/2022/11/AAPOR-Task-Force-on-2020-Pre-Election-Polling_Report-FNL.pdf (Accessed: 9 September 2026).

American Association for Public Opinion Research (2025) Task Force on 2024 Pre-Election Polling: an evaluation of the 2024 general election polls. Chaired by J. Pasek. Alexandria, VA: AAPOR, 29 October. Available at: https://aapor.org/announcements/2024-pre-election-polling-report/ (Accessed: 9 September 2026).

Ansolabehere, S. et al. (2024) 'Forecasting turnout', Harvard Data Science Review, 6(4). Available at: https://doi.org/10.1162/99608f92.62881547 (Accessed: 9 September 2026).

Keeter, S. and Igielnik, R. (2016) Can likely voter models be improved? Evidence from the 2014 U.S. House elections. Washington, DC: Pew Research Center, 7 January. Available at: https://www.pewresearch.org/methods/2016/01/07/can-likely-voter-models-be-improved/ (Accessed: 9 September 2026).

Kennedy, C. et al. (2018) 'An evaluation of the 2016 election polls in the United States', Public Opinion Quarterly, 82(1), pp. 1–33. Available at: https://doi.org/10.1093/poq/nfx047 (Accessed: 9 September 2026).

Kim, S.-y.S. and Fraga, B. (2022) When do voter files accurately measure turnout? How transitory voter file snapshots impact research and representation. APSA Preprints [preprint], 14 September. Available at: https://doi.org/10.33774/apsa-2022-qr0gd (Accessed: 9 September 2026).

More in Common (2024) Our 2024 election polling: lessons learned. London: More in Common. Available at: https://www.moreincommon.org.uk/research/our-2024-election-polling-lessons-learned/ (Accessed: 9 September 2026).

Perry, P. (1973) 'A comparison of the voting preferences of likely voters and likely nonvoters', Public Opinion Quarterly, 37(1), p. 99. Available at: https://doi.org/10.1086/268063 (Accessed: 9 September 2026).

Pew Research Center (2018) Commercial voter files and the study of U.S. politics. Washington, DC: Pew Research Center, 15 February. Available at: https://www.pewresearch.org/methods/2018/02/15/commercial-voter-files-and-the-study-of-u-s-politics/ (Accessed: 9 September 2026).

Rentsch, A., Schaffner, B.F. and Gross, J.H. (2019) 'The elusive likely voter', Public Opinion Quarterly, 83(4), pp. 782–804. Available at: https://doi.org/10.1093/poq/nfz052 (Accessed: 9 September 2026).

Explore the idea

Let’s talk

Invisible forces shape your world — until you hire Latenta®

Contact