Cognitive Abilities - What intelligence measures, and what it misses
Article L2-06
Of all psychology's ideas, intelligence is among the best measured and the most abused. There is a real, sturdy finding underneath it. There is also a century of people stretching that finding into things it never showed.
People who do well on one kind of mental test tend to do well on others, even very different ones, and that web of positive correlations is the most replicated finding in the field. We summarise it as general intelligence, or g (Spearman, 1904), and tests of it are reliable, stable, and predict real outcomes like school results, work, and health, modestly to moderately (Deary et al., 2007; Plomin and Deary, 2015). That much is solid. The overreach starts straight after. The famous "best predictor of job performance" figure was recently cut once the statistics were re-checked (Sackett et al., 2022); measured intelligence rose for most of the 20th century and has since fallen in places, both for environmental reasons, so it is clearly not a fixed innate number (Flynn, 1987; Bratsberg and Rogeberg, 2018); and intelligence being substantially heritable does not make it unchangeable or explain gaps between groups (Nisbett et al., 2012). A real, narrow, partly movable signal, routinely treated as a verdict on a person's worth.
What the science says
Consensus
Start with the one fact everything else is built on. If you give people a pile of different mental tasks, vocabulary, number patterns, mentally rotating a shape, remembering a sequence, the scores all correlate positively. Someone good at one tends to be good at the others, even when the tasks look unrelated. Charles Spearman named this the positive manifold and proposed that a single general factor sat underneath it, which he called g (Spearman, 1904). The positive manifold is about as solid as findings in psychology get, replicated across more than a century, countless test batteries, and every population studied.
On top of that base sits a tidy structure. Raymond Cattell split ability into fluid intelligence, reasoning your way through a novel problem, and crystallised intelligence, the knowledge you have accumulated (Cattell, 1963). John Carroll's survey of more than 460 datasets pulled the whole field into a three-layer map, since merged into what is now the consensus model: g at the top, around ten broad abilities beneath it (fluid reasoning, crystallised knowledge, memory, processing speed, and so on), and dozens of narrow skills below those (Carroll, 1993). One general factor, several broad ones, many specific ones.
These tests do real work. Scores are reliable and strikingly stable across a lifetime, and they predict outcomes that matter: childhood ability tracks national exam results at correlations around .5 to .8 (Deary et al., 2007), and across the lifespan measured intelligence is linked to education, occupation, health, and even how long people live (Plomin and Deary, 2015). Not perfectly, not for any one individual, but reliably on average. This is the genuine, hard-won core. Now the trouble.
Controversies
The first crack runs through the most-quoted statistic in the field. For decades the headline was that cognitive ability is the single best predictor of job performance, a correlation around .5, from a landmark meta-analysis (Schmidt and Hunter, 1998). In 2022 a team rechecked the statistical corrections that produced those numbers and found them systematically overdone: a standard adjustment for "range restriction" had been inflating the validity of cognitive tests for years (Sackett et al., 2022). Corrected, the estimates fell by .10 to .20, and cognitive ability lost its top spot to the structured interview. It still predicts performance. It is simply not the juggernaut a generation of textbooks claimed, and the episode is a clean example of a field correcting its own celebrated number.
The second crack is the Flynn effect. Across the 20th century, raw IQ scores rose by roughly three points a decade in country after country (Flynn, 1987). Genes do not change that fast, so something environmental, better nutrition, schooling, more abstract daily demands, was clearly lifting measured intelligence. And the trend is not one-way: in several wealthy countries scores have stalled or slipped in recent decades, and a study tracking Norwegian brothers showed the rise and the reversal happening within families, which pins both on environment rather than genes or changing demographics (Bratsberg and Rogeberg, 2018). Whatever the tests catch, it is movable.
The third crack is the one people fight about: heritability. Twin and genomic studies agree that intelligence is substantially heritable, and, oddly, that heritability climbs with age, from around 20% in early childhood to perhaps 80% in later adulthood (Plomin and Deary, 2015). This is where misreading does real damage, so it is worth being precise. Heritable does not mean fixed: the Flynn effect moved whole populations within a generation, and adopting children from poorer into better-off homes raises their scores by 12 to 18 points (Nisbett et al., 2012). Heritability is also a within-group statistic that depends on the environment it is measured in, it is lower among children raised in poverty, where bad environments cap potential before genes get to express it (Nisbett et al., 2012). A trait can be highly heritable within groups and still be pushed around massively by circumstances.
That precision matters most for the fourth and most charged crack: group differences and bias. Average differences between groups on these tests are real and measured. Two things are then usually overclaimed. First, that the tests must be biased: in the technical sense they largely are not, a well-built test predicts later outcomes about equally well across groups, which is what "predictive bias" means. Second, and far more important, that heritability explains the gaps: it does not. Heritability within a group says nothing about what causes the average difference between groups, the gaps have narrowed over time as circumstances changed (the Black-White gap in the United States shrank by about a third of a standard deviation; Nisbett et al., 2012), and the causes remain genuinely unresolved. Given that this is the corner of the field with the longest history of pseudoscientific misuse, the honest position is narrow: real differences in scores, contested and unsettled causes, and no warrant for the hereditarian leap.
Two smaller disputes round it out. There is an old argument about what g even is, a single underlying mental engine, or a statistical shadow cast by many overlapping skills that could, in principle, grow up reinforcing one another (Nisbett et al., 2012). And the popular alternatives that promise to dethrone IQ, Gardner's multiple intelligences and the idea of visual or auditory "learning styles", have not held up to measurement; g keeps reappearing whenever the alternatives are tested properly (Gardner, 1983).
Limitations
Whatever its strengths, the test captures a real but narrow slice of the mind, and leaves out a great deal we also call being smart: wisdom, creativity, practical judgement, character. Crystallised measures carry cultural and educational fingerprints. The scores predict averages, not the person in front of you. And a test result is co-produced by motivation, anxiety, sleep, and practice, so it is a sample of behaviour on a particular day, not a readout of fixed capacity (L2-08, L0).
Open questions
What is g, physically, in the brain? Can intelligence be raised in a way that lasts, given that most training gains fade once the training stops? What actually causes the average differences between groups? And is the general factor a real entity or a useful summary of skills that travel together?
So what
The usable core: measured ability is a real, narrow, partly movable signal that predicts averages, and almost every practical mistake comes from treating it as fixed, broad, or a measure of human worth.
For companies
Cognitive ability genuinely predicts who will do well, which is why it is tempting in hiring, but the recent correction is a direct warning against leaning on it too hard: its real-world validity is lower than the famous numbers claimed, and the structured interview now outranks it (Sackett et al., 2022). The sound approach is to combine signals, a work sample, a structured interview, and an ability measure, rather than chasing "the smartest candidate" on a single test. There is also a practical fairness and legal dimension: cognitive tests show average group differences, so relying on them heavily narrows your hiring diversity and invites challenge, which is part of why the validity-versus-diversity trade-off is worth designing around rather than ignoring.
For political parties
Intelligence is one of the easiest pieces of science to weaponise, as a marker of innate worth, a justification for who deserves what, or a cudgel in arguments about groups. The research does not support those uses. The thing is narrow, it moves with environment, and its most explosive claims (that gaps are genetic and fixed) are exactly the ones the evidence does not establish. Meritocratic rhetoric in particular often rests on a fixed-IQ myth that the Flynn effect quietly refutes: if intelligence can rise three points a decade on better conditions, it was never the pure, unearned birthright the story needs it to be.
For government
The same evidence that makes intelligence sound deterministic actually argues for investment. Environments move it: nutrition, schooling, and pulling children out of poverty all raise measured ability (Flynn, 1987; Nisbett et al., 2012), which gives early-childhood and education spending a real cognitive payoff. The flip side is a caution against sorting and tracking children early on a narrow, still-moving measure, and against reading the genuine links between ability and health or lifespan as fate rather than as one more reason to fix the environments that shape both.
How to use this
Three habits. First, treat a test score as one real but narrow signal, a sample of certain skills on a given day, never a verdict on a person's worth or potential. Second, use more than one signal for any decision that matters, since the single-number approach is both weaker and less fair than it looks. Third, hold the line against the two big overreaches: that intelligence is fixed (it visibly moves), and that heritability explains gaps between groups (it does not).
Case studies
- The Flynn effect, up and back down (Flynn, 1987; Bratsberg and Rogeberg, 2018). Raw IQ scores climbed about three points a decade through the 20th century, then stalled or fell in several countries. The reversal showed up even between brothers in the same families, which means environment, not genes or demographics, drove both directions. It is the single cleanest demonstration that "IQ" is not a fixed, innate quantity. DOI 10.1073/pnas.1718793115
- A celebrated number, corrected (Schmidt and Hunter, 1998; Sackett et al., 2022). For over twenty years, cognitive ability was taught as the best predictor of job performance, on the strength of one influential meta-analysis. When a later team re-examined the statistical corrections behind it, the validity estimates dropped by .10 to .20 and the structured interview took the top place. A model of self-correcting science, and a caution against building hiring policy on a single impressive coefficient. DOI 10.1037/apl0000994
References
- Bratsberg, B. and Rogeberg, O. (2018) 'Flynn effect and its reversal are both environmentally caused', Proceedings of the National Academy of Sciences, 115(26), pp. 6674–6678. Available at: https://doi.org/10.1073/pnas.1718793115 (Accessed: 18 June 2026).
- Carroll, J.B. (1993) Human Cognitive Abilities: a survey of factor-analytic studies. Cambridge: Cambridge University Press.
- Cattell, R.B. (1963) 'Theory of fluid and crystallized intelligence: a critical experiment', Journal of Educational Psychology, 54(1), pp. 1–22. Available at: https://doi.org/10.1037/h0046743 (Accessed: 18 June 2026).
- Deary, I.J. et al. (2007) 'Intelligence and educational achievement', Intelligence, 35(1), pp. 13–21. Available at: https://doi.org/10.1016/j.intell.2006.02.001 (Accessed: 18 June 2026).
- Flynn, J.R. (1987) 'Massive IQ gains in 14 nations: what IQ tests really measure', Psychological Bulletin, 101(2), pp. 171–191. Available at: https://doi.org/10.1037/0033-2909.101.2.171 (Accessed: 18 June 2026).
- Gardner, H. (1983) Frames of Mind: the theory of multiple intelligences. New York: Basic Books.
- Nisbett, R.E. et al. (2012) 'Intelligence: new findings and theoretical developments', American Psychologist, 67(2), pp. 130–159. Available at: https://doi.org/10.1037/a0026699 (Accessed: 18 June 2026).
- Plomin, R. and Deary, I.J. (2015) 'Genetics and intelligence differences: five special findings', Molecular Psychiatry, 20(1), pp. 98–108. Available at: https://doi.org/10.1038/mp.2014.105 (Accessed: 18 June 2026).
- Sackett, P.R. et al. (2022) 'Revisiting meta-analytic estimates of validity in personnel selection: addressing systematic overcorrection for restriction of range', Journal of Applied Psychology, 107(11), pp. 2040–2068. Available at: https://doi.org/10.1037/apl0000994 (Accessed: 18 June 2026).
- Schmidt, F.L. and Hunter, J.E. (1998) 'The validity and utility of selection methods in personnel psychology: practical and theoretical implications of 85 years of research findings', Psychological Bulletin, 124(2), pp. 262–274. Available at: https://doi.org/10.1037/0033-2909.124.2.262 (Accessed: 18 June 2026).
- Spearman, C. (1904) '"General intelligence," objectively determined and measured', The American Journal of Psychology, 15(2), pp. 201–292. Available at: https://doi.org/10.2307/1412107 (Accessed: 18 June 2026).
Explore the idea
Let’s talk
Invisible forces shape your world — until you hire Latenta®
Contact