Comparing Forecasting Methods: The Poll, the Market, and the Machine

Article M5-05

Polls and prediction markets used to forecast uncertain events on their own. Language models have added a third kind of forecaster. Because these instruments differ so much, the fair test is whether their probabilities are honest, not who called the winner.

In brief

Three tools now forecast the same uncertain events: polls and surveys, prediction markets, and machine forecasters built on large language models. They work in completely different ways. The scoreboard of who called the last election is tempting, but it says almost nothing. The fair test is calibration: across a long run of predictions, when a forecaster says an outcome is 70 per cent likely, does it happen about 70 per cent of the time? On that test, the honest verdict is modest. A machine ensemble can now match a general crowd of human forecasters, and pairing a model with the crowd beats either one alone. That combination is the real hybrid. But independent, leak-resistant benchmarks still put the best machines below the top human superforecasters. The popular claim that markets beat the polls in 2024 rests on a single election and doesn't survive a close audit. The instruments are converging, but none of them has won.

What the theory says

The theory

A forecast is a probability, not a flat call. When someone says a candidate has a 70 per cent chance of winning, they aren't predicting a win. They're giving a probability, and a probability can be tested across many forecasts.

Three different tools now produce that number. They work in such different ways that asking who called the last election isn't the right test.

The first is the poll or survey. Ask a sample of people what they think or how they intend to act, and aggregate the answers. The second is the prediction market. People bet real money on an outcome, and the contract's price moves to reflect what the crowd collectively believes, because anyone who thinks the price is wrong has an incentive to trade against it. Platforms such as Polymarket, Kalshi and PredictIt run these markets (Polymarket, 2024). The third is new: a machine forecaster is a large language model, the kind behind ChatGPT, prompted to read the available evidence and output a probability on a question whose answer isn't yet known. This third type didn't exist before 2022.

A single correct call can't tell these instruments apart. A forecaster who said 90 per cent and a forecaster who said 55 per cent on a genuine coin-flip both look correct when the event happens, and neither has shown whether its numbers can be trusted. The comparison that works is calibration. Take everything a forecaster labelled 70 per cent likely, over a long run of predictions, and check whether about 70 per cent of those things happened. A forecaster is well calibrated when its stated confidence matches reality across the board, and badly calibrated when it says 90 per cent and is right only half the time.

The tool for scoring this is old. The Brier score, introduced by Brier (1950) for weather forecasting, measures the average gap between a forecaster's probabilities and what actually happened. Lower is better. Reliability curves plot stated confidence against real frequency. None of this is new, and none of it is a governed standard. There is still no accreditation body that certifies a forecaster as calibrated. It is a shared convention, the same yardstick weather forecasters and human superforecasters have used for years (Tetlock and Gardner, 2015). What is new is that the convention now has to compare three tools instead of two.

The hybrid remains the most striking early result. In Science Advances, an ensemble of twelve language models ran against a crowd of 925 human forecasters on 31 future questions over three months, and reached accuracy that couldn't be statistically distinguished from the human crowd (Schoenegger et al., 2024). The more useful finding came next: shown the crowd's median forecast, the models improved by between 17 and 28 per cent. The machine and the crowd together beat either on its own. That combination is still the strongest single piece of evidence in this field.

A parallel result points the same way. A system built by Halawi et al. (2024) and presented at the NeurIPS conference retrieves and reads evidence before it answers, and has reported accuracy that nears an aggregated human crowd and in some settings edges past it. The team's exact scores are best treated as indicative, because they haven't yet been confirmed in the published record, but the direction is clear and independent of Schoenegger et al.'s result.

Controversies

Two active debates run through this field. Both are about how much to trust a headline claim.

The first argument is whether prediction markets beat the polls in 2024. After that US election the popular story was that the markets saw the result coming while the pollsters hedged. There is now real evidence for it. In a comparison of betting markets and polling, the markets, and Polymarket in particular, outperformed the polls, especially in the swing states that decided the outcome (Cutting et al., 2025). Their explanation leans on two things: the wisdom of a crowd with money at stake, and the difference between asking people who they intend to vote for and asking who they expect to win. But the authors are careful to note that this is one election, and a single event can't settle the question.

Set against that is the one rigorous audit, and it is far less flattering. Across more than 2,500 markets on four platforms, covering 2.4 billion dollars wagered in the final weeks of the race, the audit found a messy, inefficient market (Clinton and Huang, 2025). Accuracy depended heavily on how much money was in a given market: on some platforms most markets beat a coin flip, on others barely two-thirds did. Worse, identical contracts were priced differently on different exchanges at the same moment, a gap a truly efficient market would erase, and the mispricing got wider in the final two weeks rather than tighter. The clean story that markets outsmarted the pollsters remains mostly punditry. The closest thing to a careful measurement still says the markets were right in aggregate and sloppy in detail.

The second argument is how close machines are to humans. Author-run tournaments show the machines near level with human forecasters. Independent benchmarks, built by people with no product to sell, still put the best machines below the top humans. On ForecastBench, a benchmark designed to test models only on questions they couldn't have seen during training, the best AI system scored a Brier of 0.111 on the 2024 questions, behind the general public at 0.107 and well behind superforecasters at 0.093 (Karger et al., 2024). Lower is better, so the machine came last. Metaculus runs a live quarterly contest, its AI Benchmark Tournament, that pits forecasting bots against hand-picked human professionals on the same unresolved questions. Across the quarters from 2024 into 2025 the human pros have beaten the top bots every time, and Metaculus has reported no clear trend toward the gap closing (Metaculus, 2024). Newer systems now claim superforecaster parity on the public leaderboards, but those leaderboards move week to week, so any such figure is only as good as its date.

Limitations

There are four limitations to keep in mind.

Contamination is the first limitation. A language model is trained on a huge slice of the internet up to a cutoff date. If you test it on a question whose answer was already public before that cutoff, the model may simply be remembering the answer rather than forecasting it, which makes it look far better than it is. Paleka et al. (2025) catalogue the ways this temporal leakage inflates evaluations of machine forecasters. That work is why a serious test uses only questions that resolve after the model's training cutoff. A glowing score from a retrospective test run by the people who built the system should therefore be read with caution.

Prediction markets are still not as efficient as their supporters claim. The cross-exchange price gaps found by Clinton and Huang (2025) directly measure inefficiency: money was left on the table for weeks. A market can be a useful signal and still be beatable, thinly traded, and slow to correct.

The vendor track record is the third limitation. A handful of companies now sell machine forecasting and publish their own results. FutureSearch, for instance, currently ranks at or near the top of some neutral leaderboards. That ranking is checkable. The company also reports a high annualised return from trading on its own forecasts, but no auditable ledger sits behind that claim (FutureSearch, 2026). Mantic similarly reports a strong finish in a Metaculus contest (Mantic, 2025). A leaderboard rank is evidence. A self-reported trading return is an advertisement. A vendor page proves what is claimed. It doesn't prove the claim is true.

The weakest foundation is that the entire markets-beat-polls case, as Cutting et al. (2025) themselves point out, still rests on the 2024 US election, one event in one country. Elections are rare, expensive to forecast, and idiosyncratic. A single election is an anecdote rather than a track record.

Open questions

No one has yet independently audited whether any commercial machine forecaster is actually calibrated. So far, the vendors' strongest checkable claims are still leaderboard ranks, and a leaderboard rank is not a calibration audit. An outsider can't currently verify whether the system that finishes near the top is well calibrated or merely lucky over a short run. Kalshibench, a new benchmark, tries to score whether models know what they do not know by testing their probabilities against market prices (Nel, 2025), but it is single-authored and still barely cited. Kalshibench is an early probe rather than an answer.

It's still not clear whether the human-machine gap is closing for good. The leaderboards are still moving, agentic systems keep posting better numbers, and the marketing says the machines are pulling level. Against that, Metaculus's own tournament has so far reported no clear trend (Metaculus, 2024). The trend is noisy, and the question isn't decided yet.

There is still no governed standard for what counts as a calibrated forecaster. Brier scores and reliability curves are a shared convention. No body administers a certification, so a vendor can call its system calibrated and no one says otherwise.

Independent replication is still thin as well. The finding that a machine ensemble rivals a human crowd still rests heavily on Schoenegger et al. (2024). A recent study by Jeddi, Segovia-Martín and Servan-Schreiber (2026) points the same way, but it carries very few citations so far and does not amount to a replication base. Schoenegger's paper has drawn no published challenge so far. In a literature this young, that means the challenge hasn't been written yet, not that the result is confirmed.

The last gap is one this article searched for and still didn't find. Nobody has yet directly measured whether a poll, a market, or a machine drifts out of calibration as the world moves away from the conditions it was tuned on. A calibration built on one election year need not hold in the next. The question of forecasting drift is taken up in [M5-04], but as a framed argument rather than a settled finding, because a targeted search has so far returned nothing that measures it head-on.

So what

You now have three ways to get a probability on an uncertain event. The useful skill is reading all three honestly. The machine forecaster is now a real instrument: it can match a general crowd and improve one when paired with it, but it still can't beat the best humans on a clean test. The prediction market is a genuine but inefficient signal. The poll still measures something the other two don't: what people say about themselves. Treat the three as complementary readings. Judge a forecaster by its calibration over a run of predictions, and be most suspicious of the number that arrives with the best story attached.

For research practice

The clearest lesson in the evidence is that a machine and a crowd together beat either alone (Schoenegger et al., 2024), so the practical move is to combine them. When you evaluate any forecaster, your own or a vendor's, insist on calibration over a run of questions and refuse to be impressed by a single dramatic call. Ask when the questions resolved relative to the model's training cutoff, because a result from questions the model could have memorised is not a forecast at all (Paleka et al., 2025). Prefer the independent, forward-looking benchmark to the retrospective one the builder ran on itself. Keep the human superforecaster in the picture as the standard to beat, since on the cleanest tests that is still where the frontier sits (Karger et al., 2024). The question that matters is how often its 70 per cent comes true.

For companies

If you're buying a forecasting service, the most useful habit is to separate the checkable claim from the advertisement. A vendor's rank on a neutral leaderboard such as ForecastBench or Metaculus's FutureEval is something you can verify (Metaculus, 2025). A vendor's reported trading return, or its own retrospective accuracy, usually isn't, because no audited ledger sits behind it (FutureSearch, 2026). Ask for calibration on questions that resolved after the model's training cutoff, and ask to see the misses as well as the hits. The deeper point is that a forecast is bought to inform a real decision, so its value is in being right about the probabilities you will act on. A supplier who shows you a clean reliability curve over hundreds of resolved questions is offering intelligence. A supplier who shows only a highlight reel of correct calls is offering a story. That story is worth nothing when the outcome lands the other way and a decision was based on it.

For political parties

For political parties, the temptation is to manufacture the appearance of a forecast. A thinly traded prediction market can be moved by a single motivated bettor, and a favourable market price or a confident model output can then be publicised to create a sense of inevitability around a candidate. That is a manipulation of perception. A campaign should refuse to run it and should spot an opponent who does. A forecast is worth having only if it tells you something true about your position. You want it calibrated and honest even when it brings bad news, because a model tuned to reassure you is worse than useless in a close race. Use markets and machine forecasts to widen and speed up your reading of a volatile field, then remember that the whole case for markets beating polls rests on a single election (Cutting et al., 2025). The poll still tells you what your own voters say, which no market price can. Treat the three as cross-checks on each other, and be most careful with the one whose number you'd most like to be true.

For government and policy

For government and policy, public bodies increasingly want forecasts of disease spread, economic turns, and geopolitical events. The machine forecaster is attractive because it's fast and cheap. The responsible position is to use it as one input under a calibration standard. Before acting on any automated forecast, an agency should know whether the system has been tested on questions that resolve in the future rather than the past (Paleka et al., 2025), and whether an independent party, not the vendor, has checked its calibration. That independent audit does not yet exist for the commercial systems. That is a reason for caution. Prediction markets carry a second problem for government: they can be inefficient and open to manipulation (Clinton and Huang, 2025), so a market price is a signal to weigh. The defensible pattern is the one a careful statistical agency already uses for any model. Document what was tested and how, keep the human forecaster and the poll in the mix as cross-checks, and record which instrument informed which decision, so that when an outcome lands the forecast can be scored honestly.

How to use this

Three questions cut through almost every forecasting claim you'll meet. First, is this number calibrated over a run of predictions, or is it a single call dressed up as a track record? A forecaster who was right once is a forecaster you know nothing about. Second, who checked it, and could the questions have leaked? A result the builder ran on itself, on questions from before the model's training cutoff, is the least trustworthy kind there is. Third, what is the strongest independent evidence, and what is it worth? The peer-reviewed finding that machines rival a general crowd is real (Schoenegger et al., 2024). The claim that they beat the best humans is not yet supported on clean tests (Karger et al., 2024), and the claim that markets beat the polls rests on one election (Cutting et al., 2025). Hold each instrument to what its evidence actually shows. The machines are still approaching the best humans. The markets are contested. An honest forecaster gives you the probability that hurts and shows you the record that lets you check it.

Case studies

Metaculus AI Benchmark Tournament (2024 to 2025). This is the cleanest ongoing test of the machine-versus-human question. Each quarter, forecasting bots and a set of hand-picked human professional forecasters answer the same live, unresolved questions, so neither side can have seen the answers in advance. Across the quarters from 2024 into 2025, the human professionals have beaten the top bots every time, and Metaculus has reported no clear trend toward the machines closing the gap (Metaculus, 2024). Nobody running it has a system to sell, so it is the most useful counterweight to author-run studies that show parity. It shows where the machine still loses.

The 2.4 billion dollar election audit (Clinton and Huang, 2025). In the final weeks of the 2024 US election, 2.4 billion dollars changed hands across prediction markets. It was the largest natural test yet of whether those markets are accurate. Across more than 2,500 markets on four platforms, the markets were right in aggregate but inefficient in detail (Clinton and Huang, 2025). Accuracy rose with the money in a market. Identical contracts were priced differently on different exchanges, and the mispricing widened rather than narrowed as election day approached. The audit remains the correction to the confident post-election claim that the markets had simply outsmarted the pollsters.

Vendor claim against neutral scoreboard. A cluster of startups now sells machine forecasting. Reading their claims shows the same gap between checkable and uncheckable evidence. FutureSearch and Mantic both publish strong results. Their placements on neutral evaluations can be checked (FutureSearch, 2026; Mantic, 2025). Their headline trading returns still can't be checked, because no audited ledger backs them. The real claim and the advertised claim sit side by side on the same page, and telling them apart is the whole skill.

References

Brier, G.W. (1950) 'Verification of forecasts expressed in terms of probability', Monthly Weather Review, 78(1), pp. 1–3. Available at: https://doi.org/10.1175/1520-0493(1950)078<0001:VOFEIT>2.0.CO;2 (Accessed: 18 August 2026).

Clinton, J.D. and Huang, T. (2025) Prediction markets? The accuracy and efficiency of $2.4 billion in the 2024 presidential election. SocArXiv preprint d5yx2. Available at: https://doi.org/10.31235/osf.io/d5yx2 (Accessed: 18 August 2026).

Cutting, L.E., Hughes-Berheim, S.S., Johnson, P.M., Baroud, H. and Goldstein, B.T. (2025) Are betting markets better than polling in predicting political elections? arXiv:2507.08921. Available at: https://arxiv.org/abs/2507.08921 (Accessed: 18 August 2026).

FutureSearch (2026) Track record. Available at: https://evals.futuresearch.ai (Accessed: 18 August 2026).

Halawi, D., Zhang, F., Yueh-Han, C. and Steinhardt, J. (2024) Approaching human-level forecasting with language models. arXiv:2402.18563. Presented at NeurIPS 2024. Available at: https://arxiv.org/abs/2402.18563 (Accessed: 18 August 2026).

Jeddi, Y., Segovia-Martín, J. and Servan-Schreiber, E. (2026) 'Crowdsourced versus large language models forecasting: evidence for the accuracy–correlation effect', Philosophical Transactions of the Royal Society B, 381(1948). Available at: https://doi.org/10.1098/rstb.2024.0456 (Accessed: 18 August 2026).

Karger, E., Bastani, H., Yueh-Han, C., Jacobs, Z., Halawi, D., Zhang, F. and Tetlock, P.E. (2024) ForecastBench: a dynamic benchmark of AI forecasting capabilities. arXiv:2409.19839. Presented at ICLR 2025. Available at: https://arxiv.org/abs/2409.19839 (Accessed: 18 August 2026).

Mantic (2025) Mantic. Available at: https://www.mantic.com (Accessed: 18 August 2026).

Metaculus (2024) AI Benchmarking Tournament (AIB), Q3–Q4 2024. Available at: https://www.metaculus.com/aib/2024/q4/ (Accessed: 18 August 2026). Q3 write-up: https://www.metaculus.com/notebooks/28784/aibq3results/.

Metaculus (2025) FutureEval methodology. Available at: https://www.metaculus.com/futureeval/methodology/ (Accessed: 18 August 2026).

Nel, L. (2025) Do large language models know what they don't know? Kalshibench: a new benchmark for evaluating epistemic calibration via prediction markets. arXiv:2512.16030. Available at: https://arxiv.org/abs/2512.16030 (Accessed: 18 August 2026).

Paleka, D., Goel, S., Geiping, J. and Tramèr, F. (2025) Pitfalls in evaluating language model forecasters. arXiv:2506.00723. Available at: https://arxiv.org/abs/2506.00723 (Accessed: 18 August 2026).

Polymarket (2024) Polymarket. Available at: https://polymarket.com (Accessed: 18 August 2026).

Schoenegger, P., Tuminauskaite, I., Park, P.S., Bastos, R.V.S. and Tetlock, P.E. (2024) 'Wisdom of the silicon crowd: LLM ensemble prediction capabilities rival human crowd accuracy', Science Advances, 10(45). Available at: https://doi.org/10.1126/sciadv.adp1528 (Accessed: 18 August 2026).

Tetlock, P.E. and Gardner, D. (2015) Superforecasting: The Art and Science of Prediction. London: Random House Business Books. ISBN 978-1-84794-715-4. Available at: https://openlibrary.org/search?q=superforecasting+tetlock (Accessed: 18 August 2026).

Explore the idea

Let’s talk

Invisible forces shape your world — until you hire Latenta®

Contact