Representative Research Samples: A Representative Sample Is Now a Claim, Not a Draw

Article M0-06

Response rates collapsed, probability sampling quietly ended, and being representative became something a model asserts rather than something the fieldwork guarantees. Here is what an online sample of two thousand actually buys in 2026, and where it breaks.

In brief

For most of the last century, a survey was representative because of how people were selected: everyone in the population had a known chance of being asked. That machine has broken. Response rates fell to single digits, the industry moved online to non-probability panels, and the vast majority of survey research now recruits people who selected themselves. A sample like that is not representative by design. It is made representative afterwards, by weighting it to match the population on paper. That works, up to a point, and the point is closer than most buyers think. The same raw answers, weighted by different but defensible choices, can produce materially different results, and the errors concentrate in exactly the small groups a study is usually run to understand.

What the theory says

The theory

A probability sample earns the word "representative" from its design. If every adult has a known, non-zero chance of being selected, then statistical theory lets you generalise from the few thousand people who answered to the millions who did not. The justification lives in the sampling method, not in the data. That is the idea survey research was built on for a century, and national statistical offices still run their major surveys this way.

Almost nobody else does. Telephone response rates in the United States fell from about 36% in 1997 to 9% by 2012 and to 6% by 2018 (Pew Research Center, 2019). When only one household in twenty picks up, the known chance of selection that justified the whole approach stops describing reality, and the cost of chasing the other nineteen becomes impossible to defend. So the field moved. Between 2016 and 2022, most national pollsters in the US changed their methods, live-telephone-only polling fell to around one in ten, and online opt-in polling expanded to become the common case (Pew Research Center, 2023a). Online panel research recruited from people who volunteer, with no known probability of selection, now "constitutes the vast majority of global survey research" (Kocar and Lavrakas, 2025). Graham Kalton, tracing the century-long argument between random selection and its alternatives, puts the recent tilt toward non-probability sampling down to "the growing imperfections and costs in applying probability sample designs" alongside a new "ability to apply advanced methods for calibrating nonprobability samples to conform to external population controls" (Kalton, 2023).

And that's the pivot. When you cannot select representatively, you collect whoever you can and then adjust the result to look representative. The adjustment is weighting: you compare your sample to known population totals, from the census and other high-quality sources, and give each respondent a multiplier so the weighted sample matches the country on age, sex, region, education, and often more. A group that came in short gets weighted up. A group that came in heavy gets weighted down.

Here is the consequence that a buyer needs to hold onto. For a non-probability sample, representativeness is no longer a property of how the data was gathered. It is a property of the model you applied afterwards. As the largest review of the evidence puts it, "inferences based on nonprobability sampling are entirely dependent on models for validity" (Cornesse et al., 2020). A decade ago the profession's own task force reviewed these sources and recommended against using opt-in samples for population estimates at all, and against calling them representative (Baker et al., 2013). The field went ahead and used them anyway, because the alternative had become unaffordable. So the working reality of 2026 is a method the standards body warned against, made respectable by a modelling step that carries all the weight.

Controversies

The obvious question is whether the modelling step works. Does weighting actually recover a representative answer from an unrepresentative sample? The honest answer is: sometimes, and you cannot count on it.

The most careful recent test compared eight German probability and non-probability surveys against population benchmarks from the census. In straightforward estimates, "nonprobability surveys were clearly less accurate than probability surveys when compared with the population benchmarks" (Rohr, Silber and Felderer, 2024). The gap narrowed for relationships between variables, which is some comfort if a relationship is what you are after, but the reassurance came with a sting: the accuracy "largely depended on the variables included in the estimation." In other words, whether the adjusted answer was any good depended on the modelling choices the analyst made. The method does not deliver one representative answer. It delivers the answer implied by the model you built, and a different defensible model gives a different answer.

That is not a hypothetical. In 2016 the New York Times gave the same raw batch of Florida survey responses to four respected pollsters and asked each to weight it. They returned Clinton ahead by 4, by 3, by 1, and Trump ahead by 1, a five-point spread with no new data at all, just different reasonable choices about who to weight toward (Cohn, 2016). The pattern did not go away. Taking a 2024 survey and walking it through a sequence of individually defensible weighting decisions moved the headline margin from Harris ahead by about one point to Harris ahead by nine (Clinton, 2026). Which decisions matter is itself well studied: weighting an opt-in sample on demographics alone barely moves the estimates, while weighting on political and engagement variables moves them a lot (Pew Research Center, 2018). The person choosing the weighting variables is, quietly, choosing a large part of the answer.

None of this makes weighting worthless, and the fair counterweight matters. When a non-probability sample of social-media users was carefully adjusted with a propensity model built from demographics, technology use and political ideology, it reconciled with a parallel probability sample "for 25 of the 27 attitudes assessed" (Pollard, Robbins and Griswold, 2026). Adjustment can genuinely work. But notice what that success required: a rich set of the right variables, a probability sample to check against, and a careful academic team reporting the two attitudes it failed on as well as the twenty-five it fixed. That is the conditions under which the model earns its keep, and it is a long way from a panel dashboard producing a crosstab in an afternoon.

Limitations

The deeper problem is where the errors live, because they do not spread evenly.

When Pew compared opt-in online samples to its probability panel, the opt-in samples were off by an average of 5.8 percentage points against known benchmarks, versus 2.6 for the probability samples, roughly twice the error (Pew Research Center, 2023b). The average hides the real damage. The error reached 11.2 points for adults under 30 and 10.8 points for Hispanic adults, about four times the error for the easiest-to-measure groups. The cause was not random noise. It was a specific kind of bad respondent: people who select an agreeable answer to everything, who made up around 15% of young opt-in respondents and 19% of Hispanic ones, against 1 to 2% in the probability panel. A group that is under 1% of the population can end up as one in five of a small subsample. The survey literature has a name for the underlying behaviour, careless responding, and it is understood to produce "raw data that may not accurately reflect respondents' true levels of the constructs being measured" (Ward and Meade, 2023).

Now put that together with weighting, and you get the trap at the centre of this article. Weighting works by boosting groups that came in short, and hard-to-reach groups almost always come in short. So the same adjustment that is supposed to fix representativeness takes a scarce, contaminated cell and multiplies its influence. If a fifth of your young respondents are giving junk, and you then weight your young respondents up to their proper share of the population, you have amplified the junk, not corrected it. As far as this research could establish, nobody has published a clean measurement of how large that amplification is, so it should be read as a consequence of how weighting works rather than a number you can cite. The direction, though, is not in doubt.

Two further limits belong here. Whether the people in an online sample are even human is now a live question, because synthetic and fraudulent respondents concentrate in the same cheap, high-volume panels, and every weight you compute is calculated on a denominator you cannot fully verify. The detection side of that problem, how you tell a real respondent from a fabricated one, is a subject of its own, covered in the authenticity crisis. And the sophisticated fix is not a free pass either. Multilevel regression with poststratification, the model-based method increasingly offered as the grown-up alternative to raw weighting, can produce adjustments whose balance on the very variables you care about is worse than plain weighting's (Giordano et al., 2026). How that method works, and where it helps, belongs to small-area estimation. The point for a buyer is narrow: no weighting scheme is automatically the right one, not even the expensive one.

There is also no single thing called an online sample. Quality varies sharply by where the panel was sourced. In a head-to-head test of five common platforms, respondents on some were far more likely to pass attention checks, follow instructions and carry a unique location than respondents on others (Douglas, Ewell and Brauer, 2023). So n=2,000 online names a price, not a standard, and two samples of that size can differ in quality before a single weight is applied.

Open questions

Several things are genuinely unsettled, and saying so is part of an honest read.

The size of the amplification effect above is unmeasured. So is the share of a typical commercial panel that is now synthetic rather than human. Neither absence is a reason to ignore the direction of travel, but both are reasons not to put a number on it.

There is also no agreed way to report any of this to a buyer. The standards body has argued that the field needs quality metrics that capture "underlying representativeness" rather than the old response-rate and completion-rate numbers, and has flagged that there is no accepted method for judging a panel before you buy from it (American Association for Public Opinion Research, 2023). The metric a buyer should demand does not yet exist in a settled form.

And the largest question is left open by the polls themselves. In 2024, US election polls were more accurate than in 2016 or 2020, yet still underestimated Republican support for a third cycle running, and the firms that weighted on party identification did better than those that did not (American Association for Public Opinion Research, 2025). Weighting on the right variable helped. The catch is that nobody could be sure in advance which variable was the right one, which is exactly the problem this whole section describes.

So what

Representation used to be settled at the sampling stage and then largely forgotten. Now it is settled at the modelling stage, by choices that rarely reach the person acting on the number. Every consequence below follows from moving that decision, and from the fact that whoever makes it can move the answer.

For research practice

This is a manipulation-capable method, so start with the uncomfortable part. If different defensible weighting schemes produce different answers, then an analyst who wants a particular answer can often find a defensible route to it, by choosing which variables to weight on, which population targets to weight to, and where to trim. That is not a fringe abuse. It is the ordinary machinery of the method pointed the wrong way.

Your job is to be accountable for a number, not to tune a model until it agrees with the brief. The test to apply to any weighting decision is whether it serves the decision being made or the conclusion somebody already wanted. A crosstab engineered to support a launch is worthless as intelligence, because you have paid to hear your own hypothesis read back to you, and the market or the electorate will correct you later at full cost. The methodological discipline and the ethical one are the same move: pre-specify the weighting scheme before you see the results, write down which variables and targets you will use and why, and report the estimate under an alternative reasonable scheme so the reader can see how much of the finding is data and how much is choice. If that sensitivity is large, that is the finding.

For companies

The practical warning is about where you read the data, not whether you use it. A well-run online sample of two thousand can give you a defensible national headline. The same sample's breakdown by age band, by ethnicity, by any niche segment, is where the error was shown to be three and four times larger, and those breakdowns are usually the reason the study was commissioned. The early-adopter segment, the under-35 switcher, the high-value niche: those are small cells, and small cells are where contamination concentrates and where weighting then amplifies it.

So change what you ask a supplier. Not just how big is the sample but which variables did you weight on, to what population targets, and how much does the subgroup result move if you weight differently. Ask what share of the subgroup you are relying on came from the panel versus was manufactured by the weights. Treat nationally representative as the beginning of a question, since for a non-probability panel it is a modelling claim you are entitled to inspect. Where a subgroup is genuinely load-bearing and genuinely thin, boosting it with a real, targeted sample beats trusting a weighted handful, and the case for and against filling thin cells synthetically is its own decision, covered in validity for synthetic evidence. The blunt version: buy the headline with some confidence, and buy the crosstab with your eyes open.

For political parties

This is where the stakes and the temptation are both highest, so the ethical line comes first. The public demonstrations that a weighting choice can swing a result by several points are all political, because that is where the same raw data gets weighted in public by rival hands (Cohn, 2016; Clinton, 2026). Inside a campaign, nobody is watching, and the incentive is to keep adjusting until the number agrees with the strategy the room already committed to. A poll that reads well because you weighted it to read well has told you nothing except that you can weight a poll.

The 2024 cycle is the cautionary case in both directions. Polls were more accurate than before, and the firms that weighted on party identification were the more accurate ones (American Association for Public Opinion Research, 2025). Weighting on the right variable helped. But the pollsters who got it wrong were also making defensible choices at the time, and they only look wrong in hindsight. The lesson is not weight on party. It is that your margin is partly a modelling artefact, and you should demand to see it under more than one reasonable scheme before you bet a message or a budget on it. Defensively, the same knowledge lets you read an opponent's convenient poll: ask who was weighted up to what, and treat reluctance to say as the answer.

For government and policy

Government sits on both sides of this. Its own statistical agencies still hold the probability-sampling line for official statistics, and that discipline is worth defending precisely because so much else has abandoned it. The exposure comes from everything around the official numbers: the commissioned public-opinion research, the consultation, the quick online poll cited in a submission, most of which now runs on non-probability panels and inherits everything above.

Two levers help. The first is procurement. Public bodies buy a great deal of survey research against specifications that still ask for a sample size and a set of quality checks built for an older method. Those specifications can instead require the things that now matter: disclosure of the weighting scheme, the population targets used, reported accuracy against a benchmark, and representativeness metrics of the kind the standards body has been trying to define rather than a bare response rate (American Association for Public Opinion Research, 2023). The second is interpretation. When a result is reported for a small or hard-to-reach population, a minority community, young people, a specific region, that is exactly the estimate the evidence says is least trustworthy from a weighted online sample, and it should carry that caveat into the decision rather than being read at the same confidence as the national figure.

How to use this

Four questions, for a supplier's proposal, a colleague's deck, or a poll in the news.

  1. How were these people recruited, and does "representative" mean the design or the weighting? For almost any online sample it means the weighting, which means the next three questions apply.
  2. Which variables was this weighted on, and to what population targets? If nobody can say, nobody can defend the number, and you are looking at one modelling choice presented as the truth.
  3. How much does the answer move under a different reasonable weighting? A result that survives several defensible schemes is solid. A result that needs one specific scheme is an argument, not a measurement.
  4. Is this a headline or a subgroup? Trust the national number more than the crosstab, and trust the crosstab for a small or hard-to-reach group least of all.

A finding that holds up under all four is worth acting on. The mistake is not using weighted online data, which is most of what exists now and is often good enough. The mistake is reading a modelled number as if the model were not there.

Case studies

The New York Times Upshot (Cohn, 2016). Four professional pollsters were given the same raw Florida survey file and asked to produce an estimate. They returned four different answers spanning five points, from Clinton +4 to Trump +1, purely from different weighting and turnout choices. It remains the cleanest public demonstration that, once a sample is adjusted rather than drawn, the result is partly the analyst's decision. Everything in this article is a consequence of that one exercise.

Good Authority (Clinton, 2026). The political scientist Josh Clinton took a single 2024 survey and showed that a sequence of defensible weighting decisions moved the headline margin by roughly eight points, from a near-tie to a comfortable lead. It is the post-2022 proof that the 2016 demonstration was not a one-off curiosity of a bad year, but a permanent feature of a field that now adjusts its way to representativeness.

AAPOR 2024 pre-election polling report (American Association for Public Opinion Research, 2025). The profession's own evaluation found the 2024 polls more accurate than the two cycles before, yet still biased in the same direction for a third time, with the firms that weighted on party identification faring best. It is the honest mixed verdict: modelling can recover a good estimate, the right modelling choice matters enormously, and knowing which choice was right before the event is the part nobody has solved.

References

American Association for Public Opinion Research (2023) Data quality metrics for online samples: considerations for study design and analysis. AAPOR Task Force report, 22 February. Available at: https://aapor.org/reports/data-quality-metrics-for-online-samples-considerations-for-study-design-analysis/ (Accessed: 12 August 2026).

American Association for Public Opinion Research (2025) 2024 pre-election polling: an evaluation of the 2024 general election polls. AAPOR Task Force report, 29 October. Available at: https://aapor.org/announcements/2024-pre-election-polling-report/ (Accessed: 12 August 2026).

Baker, R., Brick, J.M., Bates, N.A., Battaglia, M., Couper, M.P., Dever, J.A., Gile, K.J. and Tourangeau, R. (2013) 'Summary report of the AAPOR task force on non-probability sampling', Journal of Survey Statistics and Methodology, 1(2), pp. 90–143. Available at: https://doi.org/10.1093/jssam/smt008 (Accessed: 12 August 2026).

Clinton, J. (2026) How pollster choices about weighting change the answer. Good Authority, 28 March. Available at: https://goodauthority.org/news/election-poll-vote2024-data-pollster-choices-weighting/ (Accessed: 12 August 2026).

Cohn, N. (2016) 'We gave four good pollsters the same raw data. They had four different answers', The New York Times (The Upshot), 20 September. Available at: https://www.nytimes.com/interactive/2016/09/20/upshot/the-error-the-polling-world-rarely-talks-about.html (Accessed: 12 August 2026).

Cornesse, C., Blom, A.G., Dutwin, D., Krosnick, J.A., De Leeuw, E.D., Legleye, S., Pasek, J., Pennay, D., Phillips, B., Sakshaug, J.W., Struminskaya, B. and Wenz, A. (2020) 'A review of conceptual approaches and empirical evidence on probability and nonprobability sample survey research', Journal of Survey Statistics and Methodology, 8(1), pp. 4–36. Available at: https://doi.org/10.1093/jssam/smz041 (Accessed: 12 August 2026).

Douglas, B.D., Ewell, P.J. and Brauer, M. (2023) 'Data quality in online human-subjects research: comparisons between MTurk, Prolific, CloudResearch, Qualtrics, and SONA', PLOS ONE, 18(3), e0279720. Available at: https://doi.org/10.1371/journal.pone.0279720 (Accessed: 12 August 2026).

Giordano, R., Cima, A., Murray, J., Hartman, E. and Feller, A. (2026) Locally equivalent weights for multilevel regression and poststratification. arXiv:2606.04250, 2 June. Available at: https://arxiv.org/abs/2606.04250 (Accessed: 12 August 2026).

Kalton, G. (2023) 'Probability vs. nonprobability sampling: from the birth of survey sampling to the present day', Statistics in Transition new series, 24(3), pp. 1–22. Available at: https://doi.org/10.59170/stattrans-2023-029 (Accessed: 12 August 2026).

Kocar, S. and Lavrakas, P.J. (2025) 'Understanding nonprobability online panel members' psychographic profiles', International Journal of Market Research, 67(5), pp. 560–588. Available at: https://doi.org/10.1177/14707853251355679 (Accessed: 12 August 2026).

Pew Research Center (2018) For weighting online opt-in samples, what matters most? 26 January. Available at: https://www.pewresearch.org/methods/2018/01/26/for-weighting-online-opt-in-samples-what-matters-most/ (Accessed: 12 August 2026).

Pew Research Center (2019) Response rates in telephone surveys have resumed their decline. 27 February. Available at: https://www.pewresearch.org/short-reads/2019/02/27/response-rates-in-telephone-surveys-have-resumed-their-decline/ (Accessed: 12 August 2026).

Pew Research Center (2023a) How public polling has changed in the 21st century. 19 April. Available at: https://www.pewresearch.org/methods/2023/04/19/how-public-polling-has-changed-in-the-21st-century/ (Accessed: 12 August 2026).

Pew Research Center (2023b) Comparing two types of online survey samples (with companion Assessing the accuracy of estimates among demographic subgroups). 7 September. Available at: https://www.pewresearch.org/methods/2023/09/07/comparing-two-types-of-online-survey-samples/ (Accessed: 12 August 2026).

Pollard, M.S., Robbins, M.W. and Griswold, M. (2026) 'A demonstration of propensity-score weighting to adjust a social media nonprobability sample survey of political attitudes', Public Opinion Quarterly, 90(2), pp. 536–565. Available at: https://doi.org/10.1093/poq/nfaf071 (Accessed: 12 August 2026).

Rohr, B., Silber, H. and Felderer, B. (2024) 'Comparing the accuracy of univariate, bivariate, and multivariate estimates across probability and nonprobability surveys with population benchmarks', Sociological Methodology, 55(1), pp. 121–154. Available at: https://doi.org/10.1177/00811750241280963 (Accessed: 12 August 2026).

Ward, M.K. and Meade, A.W. (2023) 'Dealing with careless responding in survey data: prevention, identification, and recommended best practices', Annual Review of Psychology, 74(1), pp. 577–596. Available at: https://doi.org/10.1146/annurev-psych-040422-045007 (Accessed: 12 August 2026). </content>

Explore the idea

Let’s talk

Invisible forces shape your world — until you hire Latenta®

Contact