Facial Emotion Recognition: The Face Is Not a Readout

Article M3-06

Emotion AI watches a webcam and returns a confident score for what someone feels. The post-2022 science says a facial movement is not a reliable index of a felt emotion, and in 2025 the EU made inferring emotion at work and school illegal.

In brief

Emotion AI points a camera at a face, maps the movements onto a list of emotions, and reports how happy, angry or engaged the person is. The pitch is a simple number for something that used to require a focus group. The problem is that the number rests on an idea the science has been dismantling for years: that a facial movement is a readout of a felt emotion. The strongest post-2022 reviews say movement matters but does not map one-to-one onto what a person feels, so a confident emotion score is confidence in the wrong place. The capability is real in one narrow sense. Modern face-reading models score well on benchmarks, but the benchmarks are posed or labelled expressions, not felt states, so a high accuracy figure is not proof the tool reads emotion. No commercial product publishes an independent validation, and in February 2025 the EU banned emotion inference in workplaces and schools outright. The method and the ethics lead to the same conclusion.

What the theory says

The theory

Emotion AI takes several forms; the one on trial here is facial coding by webcam. A camera captures a face, software tracks the movements of the brow, eyes and mouth, and a model maps those movements onto a set of emotion labels, then returns a score for each one. The commercial promise is scale and speed. Where reading a room once took a moderator and a small panel, a model can watch thousands of faces react to an advert or a product and hand back a tidy chart of happiness, surprise, confusion and engagement.

The underlying idea is old. Most of these tools name their theoretical basis on the page: MorphCast, for instance, describes analysing facial indicators built on Paul Ekman's six core emotions plus James Russell's circumplex model. That history goes back to the view that a small set of basic emotions produce distinct, universal facial expressions, so a smile means happiness and a scowl means anger, in Tokyo as in Toledo. If that view holds, then reading emotion off a face is a measurement problem, and a good enough model solves it.

The problem is that the view was already strongly challenged before the current wave of tools arrived. Barrett et al. (2019) reviewed the evidence for inferring emotion from facial movements and found it does not support the readout picture: people move their faces in many ways for the same emotion, and the same movement can mean different things in different situations. That review is the standard the new tools are measured against, not the thing being tested. It is the baseline that post-2022 claims must improve on.

What is truly new since 2022 is that the capability became industrialised and expanded. Face-reading models now report strong benchmark scores. Kopalidis et al. (2024) survey the modern methods and datasets and document how well the systems classify expressions. Vendors have moved beyond the basic six into much finer categories: Hume AI markets more than forty emotion categories and hundreds of output metrics. The technique has also moved into applied market research. Mancini et al. (2023) ran webcam facial coding alongside eye-tracking to study how people react to in-stream video ads, which is the practice being used. Multimodal versions that fold in voice and text are surveyed by Lian et al. (2023) and built into user-experience testing frameworks by Razzaq et al. (2023). The parallel case for voice, where the same expansion is happening, is the subject of [M1-03].

Controversies

The current disagreement is not about whether a model can classify a facial expression. It clearly can. The disagreement is whether classifying the expression tells you the emotion.

The sales pitch depends on the benchmark numbers. Kopalidis et al. (2024) document high accuracy on facial expression recognition. But accuracy against what? The models are scored against posed or pre-labelled expression categories: a picture tagged angry that the model also calls angry. That measures whether the model agrees with the label. It does not measure whether the labelled face belonged to an angry person. Krumhuber et al. (2023), the strongest single post-2022 review on this point, conclude that facial movements carry real information but do not map one-to-one onto felt emotion. So a model can be excellent at the benchmark and still be wrong about what anyone feels, because the benchmark never tested felt emotion.

Behind the benchmark is the older scientific debate, which is still unresolved. One camp holds that basic emotions reliably produce their signature expressions. Durán and Fernández-Dols (2023), replying to a defence of that position, report that basic emotions do not reliably co-occur with the facial expressions predicted for them: the felt state and the expected face differ more often than the theory allows. Barrett et al. (2025) restate the constructionist account, in which an emotion is assembled in context rather than broadcast by a dedicated facial signal. The debate is genuine and two-sided, which is exactly why building a product that assumes one side has won is a risk the buyer is rarely told about.

Then there is the expansion strategy. If six emotions are crude, the reasoning goes, forty are better, so a tool that reports more categories reports something finer and more valid. That does not follow. Adding categories increases the number of readings without validating any of them, and no vendor page shows independent evidence that the expanded set reads true. Doerfler and Stark (2024), studying how emotion tracking is sold in driver-monitoring systems, show how the framing itself works: presenting the tool as a settled instrument legitimates a capability the science has not settled.

Limitations

Even granting the tools their best case, three limits are well documented.

The first is that the evidence base uses posed faces, not real faces. Dawel, Krumhuber and Palermo (2025) argue that emotion research needs spontaneous, naturalistic expressions rather than acted ones, because acted expressions are cleaner and more exaggerated than anything a real person produces watching an advert at their desk. A model trained and tested on posed faces will overstate how well it does on the distracted, half-attending, real one. The early attempts to test the tools on spontaneous rather than acted expressions are exactly the ones the vendors do not run: Donovan et al. (2024), in a preprint that has not yet been peer reviewed, put automated emotion detection to a direct validity test on spontaneous expressions and treat the result as an open question rather than a pass. Cross, Acevedo and Hunter (2023) make the applied version of the point for practitioners: automated facial-coding tools carry specific validity limits that anyone deploying them needs to know before trusting a score.

The second is culture. The readout picture assumed expressions are universal, but the field is now building culture-specific datasets precisely because a single model does not work across cultures. Mishra, Bhushan and Venkatesh (2025) introduce an Indian spontaneous micro-expression dataset, and the existence of such work is itself an admission: if one model generalised across cultures, nobody would need to collect a separate one for each. A preprint by Jamróz, Wysocka and Garbat (2026) catalogues failure modes of deployed facial emotion recognition systems, though as a preprint it has not yet been through peer review.

The third is that these systems are deployed at scale while remaining weakly validated. There is no independent, method-backed accuracy figure for any commercial product: no vendor publishes a held-out benchmark, a sensitivity and specificity pair, or a third-party audit, and the highest concrete number on offer, Entropik's claim of more than 95% accuracy across the basic six plus contempt, comes with no test set behind it. Katirai (2023) reviews the ethics literature and finds emotion recognition rolled out faster than its validity or social risks have been worked out, and Walsh (2025) asks whether a face can speak for itself. The industry's answer to the validity gap is usually size: Affectiva, now part of Smart Eye, advertises what it calls the world's largest emotion AI database, millions of face videos and billions of frames across dozens of countries. But corpus size is not sensitivity or specificity; it says how much the model has seen, not whether what it reports is true.

Open questions

Two questions are genuinely open. The first is whether adding channels helps: multimodal systems that read face, voice and text together report better numbers (Lian et al., 2023; Razzaq et al., 2023), but they too are validated against expression labels rather than felt states, so whether they read emotion better or only classify expressions better is unknown.

The second is whether the EU ban reshapes the market at all. Its prohibition targets workplaces and schools, and most commercial emotion measurement sits outside those walls, so whether enforcement reaches further, and whether the market shifts if it does not, is untested.

So what

This method is not just weak; it is intrusive by design, and the two problems are connected. A tool that infers a person's inner state from their face, usually without a clear moment of consent, is a surveillance instrument as much as a measurement one, and in February 2025 the EU treated it as such by banning emotion inference in workplaces and education under the AI Act (European Union, 2024). The methodological and ethical arguments lead to the same conclusion. The tool does not reliably detect felt emotion, and detecting felt emotion in people who did not agree to it is exactly what a regulator has now banned in its most sensitive settings. For all four audiences below, the honest approach is the same: treat an emotion score as a hypothesis about a feeling, not a reading of one.

For research practice

Treat facial coding as a signal about attention and expression, not a measurement of emotion. It can tell you a face moved when the advert changed, which is genuinely useful for spotting the moment something worked or lost people's attention. It cannot tell you, on its own, that the movement was delight rather than confusion, because the science says that movement does not map cleanly onto the feeling (Krumhuber et al., 2023). The decision test applies directly: does the emotion score serve the question you are answering, or does it just make a weak finding look more precise? If a deck reports that an advert scored 0.72 on joy, ask what that number was validated against. The published benchmarks measure agreement with posed expression labels, not felt emotion (Kopalidis et al., 2024), so a high vendor accuracy figure does not guarantee what it seems to guarantee. Pair any facial-coding read with something that reaches the felt state more directly, such as what people say and do afterwards, and let the two check each other.

For companies

The claim is that emotion AI replaces slow, expensive qualitative work with a scalable way to measure how customers feel. The honest version is narrower and still useful: it is a cheap way to see where attention moves across a lot of content. The mistake is to buy the emotion label and act on it as fact. A concept test that reports audiences felt trust does not fail clearly; it produces a confident chart based on a reading the science does not support, and the error is invisible in the results until the product ships to an audience that did not feel what the model said. When a vendor advertises a huge database or a high accuracy number, ask the two questions that separate research from performance. Is the accuracy measured against felt emotion or against labelled expressions? And is the database offered as evidence the tool reads emotion, or in place of that evidence? Corpus size is not validity (Affectiva, 2026). A tool that flatters your creative is exposed when the campaign underperforms, and you will have paid for a mirror instead of an early warning.

For political parties

Reading a crowd's or a focus group's faces to gauge reaction to a message is an old ambition, and emotion AI promises to automate it. The same caution is sharper here, because the temptation is to treat a reassuring emotion read as proof a message works. It is not. The tool reports the expressions of the people in front of the camera, and those expressions do not reliably reveal the feeling, especially across the varied faces of a real electorate, where cross-cultural performance is a known weak point (Mishra, Bhushan and Venkatesh, 2025). The rigorous approach and the responsible approach are identical. Use facial coding, if at all, to flag moments worth investigating, then test those moments on real people who tell you in their own words what they thought. A campaign that trusts the emotion score more than the voter has obtained a number that matches its own view, which has no value when the vote occurs.

For government and policy

Two things matter here, and the EU has acted on both. The AI Act now bans inferring emotion in workplaces and education, in force since 2 February 2025, with a related disclosure duty following in August 2026 (European Union, 2024). The machinery of the Act is covered in [M6-06], so the focus here is more specific: a regulator has banned the main feature that vendors sell, and it did so on grounds of both intrusion and unreliability. A standards body has reached a compatible conclusion. IEEE (2024) published a standard on ethical considerations in emulated empathy, treating emotion systems as not yet reliable and needing rules, not as a ready-to-use technology. For any public body weighing an emotion-reading tool for services or research, the safe approach is to assume the reading is not validated until an independent audit confirms it, to avoid using it for important decisions about individuals, and to require telling people whenever it is used.

How to use this

When someone offers you an emotion score from a face, run three checks before you believe it. First, ask what the accuracy figure was measured against. If the answer is posed or labelled expressions rather than felt emotion, the number tells you the model matches a human labeller, not that it read a feeling. Second, ask for independent validation, not database size. A count of faces seen is not sensitivity or specificity, and no commercial product publishes the latter. Third, ask where the tool is being used. If it infers emotion on people at work or in education inside the EU, it is now prohibited (European Union, 2024), and if it is used anywhere else, disclosure and consent are the minimum the setting demands. Facial coding earns a place as a pointer to moments worth a closer look. It has not earned the place it is sold for, which is a verdict on what a person feels.

Case studies

The European Union AI Act (2024). The clearest post-2022 verdict on emotion AI is a statutory one. Under Article 5 of the AI Act, inferring emotions from biometric data in the workplace and in education is a prohibited practice, in force from 2 February 2025, with a further transparency duty under Article 50 requiring people to be told when they are exposed to an emotion recognition system, applying from 2 August 2026 (European Union, 2024). A regulator drew the line at exactly the capability vendors market most confidently, and it drew it on both grounds traced above: the inference is intrusive, and it is not reliably valid. The Act's wider machinery is covered in [M6-06]; the point for emotion AI is that the most sensitive uses are now simply off the table in one of the world's largest markets.

Affectiva and the scale claim (2026). Affectiva, part of Smart Eye, is the clearest example of size standing in for validity. Its public materials advertise what it calls the world's largest emotion AI database, measured in millions of face videos and billions of frames gathered across dozens of countries (Affectiva, 2026). The figure is impressive and it is real, but it answers a different question than the one a buyer should be asking. It says how much the system has seen. It does not say whether what the system reports about a felt emotion is correct, and no independent accuracy figure accompanies it. This is not a failure story. It is the ordinary case, and it shows the substitution the whole trial is about: a number that measures the corpus, offered as if it measured the truth of the read.

References

Affectiva (2026) Emotion AI: Media Analytics and Emotion SDK (part of Smart Eye). Available at: https://www.affectiva.com/ (Accessed: 18 August 2026).

Barrett, L.F., Adolphs, R., Marsella, S., Martinez, A.M. and Pollak, S.D. (2019) 'Emotional Expressions Reconsidered: Challenges to Inferring Emotion From Human Facial Movements', Psychological Science in the Public Interest, 20(1), pp. 1–68. Available at: https://doi.org/10.1177/1529100619832930 (Accessed: 18 August 2026).

Barrett, L.F., Atzil, S., Bliss-Moreau, E., Chanes, L., Gendron, M., Hoemann, K., Katsumi, Y., Kleckner, I.R., Lindquist, K.A., Quigley, K.S., Satpute, A.B., Sennesh, E., Shaffer, C., Theriault, J.E., Tugade, M. and Westlin, C. (2025) 'The Theory of Constructed Emotion: More Than a Feeling', Perspectives on Psychological Science, 20(3), pp. 392–420. Available at: https://doi.org/10.1177/17456916251319045 (Accessed: 18 August 2026).

Cross, M.P., Acevedo, A.M. and Hunter, J.F. (2023) 'A Critique of Automated Approaches to Code Facial Expressions: What Do Researchers Need to Know?', Affective Science, 4(3), pp. 500–505. Available at: https://doi.org/10.1007/s42761-023-00195-0 (Accessed: 18 August 2026).

Dawel, A., Krumhuber, E.G. and Palermo, R. (2025) 'Faking It Isn't Making It: Research Needs Spontaneous and Naturalistic Facial Expressions', Affective Science, 7(1), pp. 88–103. Available at: https://doi.org/10.1007/s42761-025-00320-1 (Accessed: 18 August 2026).

Doerfler, A. and Stark, L. (2024) 'Legitimating Emotion Tracking Technologies in Driver Monitoring Systems', Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 7, pp. 396–410. Available at: https://doi.org/10.1609/aies.v7i1.31645 (Accessed: 18 August 2026).

Donovan, R., O'Reilly, R., Johnson, A. and De Roiste, A. (2024) 'Evaluating the Validity of Automated Emotion Detection from Spontaneous Expressions across Modalities: Insights from the PEM Dataset'. OSF Preprints. Available at: https://doi.org/10.31234/osf.io/skru2 (Accessed: 18 August 2026).

Durán, J.I. and Fernández-Dols, J.-M. (2023) 'Basic emotions do not reliably co-occur with predicted facial expressions: Reply to Witkower et al. (2023)', Emotion, 23(3), pp. 908–910. Available at: https://doi.org/10.1037/emo0001227 (Accessed: 18 August 2026).

Entropik (2026) Decode: Emotion AI and Insights platform. Available at: https://www.entropik.io/ (Accessed: 18 August 2026).

European Union (2024) Regulation (EU) 2024/1689 (Artificial Intelligence Act), Article 5(1)(f) and Article 50(3). Available at: https://artificialintelligenceact.eu/article/5/ (Accessed: 18 August 2026).

Hume AI (2026) Expression Measurement API. Available at: https://www.hume.ai/ (Accessed: 18 August 2026).

IEEE (2024) IEEE 7014-2024: IEEE Standard for Ethical Considerations in Emulated Empathy in Autonomous and Intelligent Systems. Available at: https://standards.ieee.org/ieee/7014/7648/ (Accessed: 18 August 2026).

Jamróz, A., Wysocka, P. and Garbat, P. (2026) 'Weaknesses of Facial Emotion Recognition Systems'. arXiv:2601.12402. Available at: https://arxiv.org/abs/2601.12402 (Accessed: 18 August 2026).

Katirai, A. (2023) 'Ethical considerations in emotion recognition technologies: a review of the literature', AI and Ethics, 4(4), pp. 927–948. Available at: https://doi.org/10.1007/s43681-023-00307-3 (Accessed: 18 August 2026).

Kopalidis, T., Solachidis, V., Vretos, N. and Daras, P. (2024) 'Advances in Facial Expression Recognition: A Survey of Methods, Benchmarks, Models, and Datasets', Information, 15(3), 135. Available at: https://doi.org/10.3390/info15030135 (Accessed: 18 August 2026).

Krumhuber, E.G., Skora, L.I., Hill, H.C.H. and Lander, K. (2023) 'The role of facial movements in emotion recognition', Nature Reviews Psychology, 2(5), pp. 283–296. Available at: https://doi.org/10.1038/s44159-023-00172-1 (Accessed: 18 August 2026).

Lian, H., Lu, C., Li, S., Zhao, Y., Tang, C. and Zong, Y. (2023) 'A Survey of Deep Learning-Based Multimodal Emotion Recognition: Speech, Text, and Face', Entropy, 25(10), 1440. Available at: https://doi.org/10.3390/e25101440 (Accessed: 18 August 2026).

Mancini, M., Cherubino, P., Martinez, A., Vozzi, A., Menicocci, S., Ferrara, S., Giorgi, A., Aricò, P., Trettel, A. and Babiloni, F. (2023) 'What Is behind In-Stream Advertising on YouTube? A Remote Neuromarketing Study employing Eye-Tracking and Facial Coding techniques', Brain Sciences, 13(10), 1481. Available at: https://doi.org/10.3390/brainsci13101481 (Accessed: 18 August 2026).

Mishra, R., Bhushan, B. and Venkatesh, K.S. (2025) 'Toward cross-cultural emotion detection: the Indian spontaneous micro-expression dataset', Frontiers in Psychology, 16. Available at: https://doi.org/10.3389/fpsyg.2025.1656104 (Accessed: 18 August 2026).

MorphCast (2026) Emotion AI Engine. Available at: https://www.morphcast.com/ (Accessed: 18 August 2026).

Razzaq, M.A., Hussain, J., Bang, J., Hua, C.-H., Satti, F.A., Rehman, U.U., Bilal, H.S.M., Kim, S.T. and Lee, S. (2023) 'A Hybrid Multimodal Emotion Recognition Framework for UX Evaluation Using Generalized Mixture Functions', Sensors, 23(9), 4373. Available at: https://doi.org/10.3390/s23094373 (Accessed: 18 August 2026).

Walsh, E. (2025) 'Does a Face Speak for Itself? Emotion Recognition Technologies and Explainable AI', Philosophy & Technology, 38(2). Available at: https://doi.org/10.1007/s13347-025-00891-8 (Accessed: 18 August 2026).

Explore the idea

Let’s talk

Invisible forces shape your world — until you hire Latenta®

Contact