When the science doesn't replicate
Lesson 0.03
Some of the most famous findings in behavioural science fell apart when they were redone. That isn't a scandal, it's quality control, and it tells you exactly how much to trust a single study.
When independent teams redid 100 published psychology studies, only about a third got the same result, and the effects that did survive were roughly half as strong (Open Science Collaboration, 2015). Beloved findings like "ego depletion" and the physical claims behind "power posing" largely vanished under careful testing. The causes are mundane: small samples, flexible analysis, and journals that reward surprising results. The takeaway is not that behavioural science is worthless. It is that any single study is provisional, and the findings worth betting on are the replicated, well-powered, preregistered ones.
What the science says
Consensus
In 2015 the Open Science Collaboration published the result that named the problem. Independent teams repeated 100 studies from three top psychology journals. About 97% of the originals had reported a clear positive result; only around 36% of the repeats did, and even the effects that held up came in at roughly half their original size. A few years later, Camerer and colleagues (2018) redid 21 social-science experiments from Nature and Science and found 13 replicated, again at about half strength. So the picture is not "nothing is real." It is "most of the careful, top-tier work holds, and a meaningful slice does not."
Some specific casualties are worth naming, because they were famous. Ego depletion, the idea that willpower runs down like a fuel tank, was a textbook effect with thousands of citations (Baumeister et al., 1998). When 23 labs ran the same preregistered test on roughly 2,000 people, the effect came out essentially zero (Hagger et al., 2016). Power posing fared a little better and a lot worse: a larger study found that standing in a confident pose made people feel more powerful but did not move their hormones or their risk-taking (Ranehill et al., 2015), and the original's first author later said publicly that she no longer believed the effect was real. Social priming results, like the classic finding that words about old age made people walk more slowly, failed to reproduce, with the experimenters' own expectations implicated (Doyen et al., 2012). The causes turned out to be ordinary: small samples, flexible analysis choices that quietly manufacture significance (Simmons, Nelson & Simonsohn, 2011), and a publishing system that prizes the surprising over the solid, exactly as Ioannidis (2005) had warned.
Controversies
How bad is it, really? That part is genuinely disputed. Gilbert and colleagues (2016) reanalysed the Open Science Collaboration data and argued that once you account for low statistical power and the fact that many repeats used different populations, the results are consistent with reproducibility being quite high. The original team pushed back, and the exchange sharpened how the field now defines a "successful" replication. So the headline number moves depending on how you count. What almost nobody disputes is the direction: a real chunk of the published literature is fragile, and the methods that produced it were too loose.
Limitations
A failed replication is not a verdict on its own. It can mean the original was a false positive, or that the new study differed in population, context, or sample size. This is why single replications rarely settle anything and multi-lab preregistered efforts carry the weight. It is also why the crisis is best understood as concentrated in specific areas, mainly social and cognitive psychology and parts of biomedicine, rather than a blanket indictment of "science."
Open questions
We still don't know the true replicability rate for most subfields, or how much of the failure reflects false originals versus effects that are real but fussy about context. And the big practical question is open: whether the reforms now spreading, preregistration, registered reports, and bigger samples, are actually raising the hit rate. The early signs are encouraging, but the measurement is ongoing.
So what
This is the discipline that earns everything else in this series its credibility. If you are going to act on a behavioural finding, you need a way to tell the durable ones from the fashionable ones. The replication crisis hands you that filter.
For companies
Be wary of strategy built on a single exciting study. The findings that made the best conference talks (power posing, subtle priming, willpower as a fuel tank) are among those that failed to hold up. Before you spend budget on a tactic, ask three questions: has it replicated, how big is the effect when it does, and was the test preregistered? A flashy result from one small study with a huge effect is a warning sign, not a green light.
For political parties
Campaign folklore is full of one-study tricks: a subliminal cue, a single clever message that supposedly swung a race. Most of these do not survive contact with a large field experiment. Trust effects that show up repeatedly, at scale, in real elections, and discount the cute lab result that a vendor is selling as a silver bullet.
For government
Policy is where fragile findings do real damage, because they get deployed at scale. Lean on replicated, field-tested evidence, expect effects to shrink when they leave the lab, and build evaluation into the rollout so a weak intervention gets caught early rather than scaled blindly.
How to use this
Treat every single study as provisional. Weight evidence by whether it has replicated, by the size and stability of the effect rather than mere statistical significance, and by whether the design was preregistered. The rule of thumb that will rarely fail you: prefer boring and robust to exciting and fragile.
Case studies
- Ego depletion (1998 to 2016). A cornerstone of self-control research, cited thousands of times, that came out essentially zero when 23 labs tested it together on around 2,000 people (Hagger et al., 2016). The clearest example of a beloved effect that the evidence simply did not support. DOI 10.1177/1745691616652873
- Power posing. A larger, tighter study found the confident pose changed how people felt but not their hormones or their choices (Ranehill et al., 2015), and the original's first author publicly stopped defending the effect. A rare and honest example of a scientist following the evidence away from her own famous result. DOI 10.1177/0956797614553946
References
- Open Science Collaboration (2015), Science. DOI 10.1126/science.aac4716
- Camerer et al. (2018), Nature Human Behaviour. DOI 10.1038/s41562-018-0399-z
- Ioannidis (2005), PLoS Medicine. DOI 10.1371/journal.pmed.0020124
- Simmons, Nelson & Simonsohn (2011), Psychological Science. DOI 10.1177/0956797611417632
- Baumeister et al. (1998), JPSP. DOI 10.1037/0022-3514.74.5.1252
- Hagger et al. (2016), Perspectives on Psychological Science. DOI 10.1177/1745691616652873
- Carney, Cuddy & Yap (2010), Psychological Science. DOI 10.1177/0956797610383437
- Ranehill et al. (2015), Psychological Science. DOI 10.1177/0956797614553946
- Gilbert et al. (2016), Science. DOI 10.1126/science.aad7243
- Doyen et al. (2012), PLoS ONE. DOI 10.1371/journal.pone.0029081
Explore the idea
Let’s talk
Invisible forces shape your world — until you hire Latenta®
Contact