Research Modelling: Causality, Segmentation and Uncertainty
Privacy regulation did what methodologists could not: it forced the industry off user-level tracking and back to aggregate causal models. Around that forced return, a new computational toolkit arrived — effects that vary by person, meaning measured in vectors, findings wired into graphs. This chapter is about the engine that turns data into answers, and the uncertainty it is tempted to hide.
M4 / 11 published articlesThe model is the new instrument
For about a decade, the dominant way to measure marketing was to follow individuals. Track a user across devices, credit the ads they saw before they bought, and assign a return on investment to each touchpoint. It was granular, it was intuitive, and it was built on a surveillance infrastructure that privacy regulation has now substantially dismantled. When Apple let users opt out of cross-app tracking in 2021, and when browsers began retiring third-party cookies, the attribution model that powered a generation of digital marketing quietly broke.
What came back was something older, rebuilt. Marketing mix modelling — the aggregate, regression-based approach that the ad industry used before it could follow individuals — returned, but in a form the original practitioners would barely recognise. The new versions are Bayesian, open-source, calibrated by experiments, and transparent enough that a client can inspect the priors rather than trusting the black box. Around that core, a broader computational revolution arrived at the same time: causal machine learning that estimates not just whether something works but for whom; small-area estimation that draws big answers from small samples; data fusion that stitches a single person's picture from scattered datasets; embeddings that turn meaning into geometry; knowledge graphs that wire findings into navigable maps.
This chapter is about that engine. It is more powerful than what it replaced, and it is more honest, but only when the people using it let the uncertainty through.
Why this matters to you
If you allocate budget against evidence — media spend, campaign choices, segmentation strategies, pricing — the models that generate your evidence have changed. The new toolkit gives you things the old one could not: uncertainty intervals instead of point estimates, heterogeneous effects instead of averages, structured knowledge instead of loose decks. But it also makes a new demand: you have to be willing to act on a range rather than a number, and the organisation has to be willing to hear "we don't know" as a legitimate finding. Every model in this chapter produces a distribution. Every dashboard is tempted to collapse it to a dot. The chapter exists to make sure you see the distribution.
What you'll find inside
The pieces here move from the anchor model outward to the frontier. You will start with the MMM renaissance and what the new open-source implementations actually deliver compared to the old proprietary ones. You will meet incrementality testing — the insistence that if you cannot show a causal lift, you have not shown anything — and the triangulated-measurement frameworks that combine models, experiments and attribution into something more trustworthy than any one alone. Then the chapter widens: causal machine learning that asks "for whom does this work?"; MRP and small-area estimation that squeeze national precision from local samples; data fusion that unifies fragmented records; embeddings as research instruments that measure meaning directly; knowledge graphs that turn a library of findings into a queryable structure. You will see choice modelling rebuilt with deep learning, and the quiet revolution in treating uncertainty as something you deliver to the client rather than something you smooth away. The chapter closes with segmentation that survives — clusters that actually replicate — and the new contamination risk when an LLM can generate the segment profiles it was asked to find.
The honest note
The characteristic caveat for this chapter is about the gap between what a model produces and what a dashboard shows. Every serious model in this chapter outputs a posterior distribution, a confidence interval, a range of plausible answers. That range is the most valuable part of the output, because it tells you how much you know and how much you are guessing. But distributions are hard to act on, hard to present, and hard to sell, so the pressure at every step is to collapse the range to a single number and move on. The result is false precision: a finding that looks certain because the uncertainty was hidden, not because it was resolved. The discipline this chapter asks for is simple and difficult — show the interval, explain what it means, and let the decision-maker decide how much risk to carry. A model that hides its own uncertainty is more dangerous than no model at all.
The eleven pieces in this chapter
- "The big model came back" — the MMM renaissance
- "Lift is the only currency" — incrementality or it didn't happen
- "Stop hunting the one true number" — triangulated measurement
- "For whom does it work?" — causal machine learning
- "Big answers from small samples" — MRP and small-area estimation
- "One person, many datasets" — data fusion
- "Meaning, measured in vectors" — embeddings as instruments
- "Findings become a map" — knowledge graphs for insight
- "Utility theory meets deep learning" — choice modelling's new engines
- "Intervals grow teeth" — uncertainty as a deliverable
- "Clusters that replicate" — segmentation that survives
Read the articles
- M4-01
Marketing Mix Modelling: The Big Model Came Back
- M4-02
Advertising Incrementality: Incrementality or It Didn't Happen
- M4-03
Triangulating Campaign Measurement: The Campaign Ran. Four Instruments Can't Agree on What It Did.
- M4-04
Uplift Modelling: Knowing Who Will Buy Is Not Knowing Whom You Moved
- M4-05
Small-Area Market Estimation: Big Answers from Small Samples
- M4-06
Data Fusion in Market Research: The Correlations Your Data Never Measured
- M4-07
Brand Positioning Maps from Text: Draw the Market Map Without Asking Anyone
- M4-08
Research Knowledge Graphs: Let a Machine Draw the Map, Then Check What It Invented
- M4-09
Choosing a Conjoint Engine: The Choice Engine That Fits Better and Reads Worse
- M4-10
Research Uncertainty and Intervals: The Number Is Only Half the Deliverable
- M4-11
Customer Segmentation Validation: Make Your Segments Prove They Are Real
Let’s talk
Invisible forces shape your world — until you hire Latenta®
Contact