Marketing Mix Modelling: The Big Model Came Back
Article M4-01
Marketing mix modelling was the old way to measure advertising, and the collapse of user-level tracking brought it back, now as open-source tools anyone can inspect. The awkward part is that the analyst's assumptions, not the data, often decide the answer.
In brief
For years the common way to measure advertising was to follow individual users and credit the ads they saw before they bought. Privacy changes made that impossible. When Apple let iPhone users switch off app tracking in 2021, and identifiers became less useful across the web afterward, tracking individual users' ad effects stopped being reliable. Measurement moved back to an older method: marketing mix modelling, which never watches individuals at all. It takes aggregate spend and sales, week by week, and estimates what each channel contributed. The method is decades old. What is new, and genuinely post-2022, is a set of open-source tools that made an inspectable version of it available to everyone, led by Google's Meridian. The honest limitation applies to the whole field. These models are often weakly identified, which means the data alone cannot determine the answer, so the analyst's starting assumptions determine the result. That makes the technique powerful and quietly dangerous: a prior can be set until a channel reports the return its owner was hoping for. The rigorous solution and the responsible solution are the same.
What the theory says
The theory
Every way of measuring advertising has to answer one question: of everything that happened, what did the advertising actually cause? For most of the 2010s the popular answer was to follow people. Track a user across sites and apps, note which ads they saw, and credit the last one, or spread credit across the path, before they bought. This is multi-touch attribution, and it depends on recognising the same person again and again.
That recognition stopped working. In 2021 Apple began asking iPhone users whether an app could track them, and most said no. The identifiers that let advertisers stitch a person's journey together decayed over the years that followed. It is worth being precise about the cause, because the popular version is wrong. This was not the death of the browser cookie. Google reversed its plan to remove third-party cookies from Chrome in July 2024. The main cause was the loss of reliable user-level identity, not any one technology (Runge et al., 2024).
When you can no longer follow the individual, you go back to the aggregate. Marketing mix modelling never watched individuals in the first place. It takes what a business already has, weekly spend on each channel and the sales that followed, and estimates how much each channel contributed. Two ideas do most of the work. Advertising continues to have an effect after it runs, so this week's sales owe something to last week's ads, an effect called adstock or carryover. And each channel eventually tires: the first money spent has more effect than the last, which is saturation, or diminishing returns. Jin et al. (2017) set out the Bayesian version of these mechanics at Google, and that pre-2022 paper is the foundation the newer tools build on.
None of that is new, and that is the point: the post-2022 development is not the model itself but the tooling around it, concrete and dateable. A class of open-source tooling now exists that did not exist in this form before. Google's Meridian reached its first stable release early in 2025 under an open licence, replacing an earlier unofficial Google library that was formally retired in January 2026 (Google, 2025; Google, 2026). Meta's Robyn and the community-run PyMC-Marketing round out the set. For the first time a practitioner can read the actual code that produces the number, rather than trusting a consultancy's black box. Open-sourcing the method has become the field's way of showing it is transparent.
One more idea sets up the rest of the article. In the Bayesian tools, Meridian and PyMC-Marketing, the analyst sets priors. A prior is a starting assumption, stated as a range, about how much a channel probably does before the data is consulted. The model combines that assumption with the data and reports what it now believes, along with how uncertain it is. Priors are not a flaw. Stating your assumptions openly and letting the data revise them is the central honest practice of the Bayesian approach (PyMC Labs, no date). But it does mean a human choice is part of the model, and the next sections are about how much that choice changes the answer.
Controversies
The first problem is with the word Bayesian, which is used to describe tools that work differently. Meridian and PyMC-Marketing are genuinely Bayesian: explicit priors, full uncertainty, all the expected features. Robyn is not. It uses ridge regression, a penalised form of ordinary regression that reduces estimates toward zero to keep them stable, tuned by an evolutionary search and a seasonality routine (Meta, no date). It has no priors in the Bayesian sense. This matters for a reader comparing tools, because saying we use Bayesian MMM does not guarantee the same thing. Robyn's equivalent of a prior is its regularisation and its hyperparameter settings, which are analyst choices too, just under different names. The three tools share being open-source, not being Bayesian.
The deeper controversy is about what the priors are really doing. One view is that a prior contains real prior knowledge and the data improves it. A stricter view has become more common. Dew, Padilla and Shchetkina (2024), in a preprint focused on the claim that the typical marketing mix model is flawed, argue that the effects these models try to estimate are frequently not identifiable from the data at all. Identifiable means the data contains enough information to determine one answer rather than many. Their argument is that the saturation curves and the way effects shift over time can take quite different shapes while fitting the same historical data equally well. The problem is worst in typical marketing data, where channels move together and spend changes smoothly. When the data cannot choose between competing answers, the prior or the regularisation chooses. So the perspective changes. It is not just that priors slightly change the result. When identification is weak, the assumption decides and the data mostly agrees.
That change in view is not a reason to abandon the method. It is a reason to control it, and everyone points to the same fix. Calibrate the model against a real experiment, usually a geographic test where spend is deliberately raised in some regions and held in others. The experiment gives you a real result to compare the priors with. The mechanics of those tests belong to geo experiments and are not re-taught here. What matters for this article is the logic: an experiment is how a prior becomes useful instead of just being assumed.
Limitations
The core limitation is the identification problem above, and it is structural rather than a flaw in any one tool. No amount of clean code fixes data that cannot distinguish between two possible explanations. This is no longer one team's claim. Alongside Dew, Padilla and Shchetkina (2024), a separate preprint by Marín (2023) reaches a related conclusion from a different starting point, that these models tend to over-credit the channels which received the most spend and struggle to pull one channel's effect apart from another's. Both are preprints rather than peer-reviewed papers, so the settled literature is still thin. The point itself is a mathematical one about what observational spend data can and cannot reveal, and it is the central and reliable fact of the topic.
A close relative is collinearity, which is the plain reason channels are hard to separate. If a brand always runs television and online video together, in the same weeks and the same proportions, the model has no way to tell which one did the work, because it never saw one without the other. Wu et al. (2026) frame this as the notorious problem it is and propose grouping regions by how their spend patterns correlate, then fitting the model at the group level. That work was presented at academic workshops in 2023 and posted to a preprint server later, so it is best read as 2023-era thinking despite the citation year. It is a proposed fix, not a validated standard.
Then there is the nature of the data itself. Marketing mix models usually run on weekly figures, which means a couple of hundred observations at most, against which the model tries to estimate carryover and saturation shapes for every channel. That is many estimates to make from few data points, and it is precisely why the priors and the regularisation become essential. The less data there is, the more the assumptions matter. There is no clean public benchmark settling how short a series is too short, so this is a structural caution rather than a measured threshold.
The most serious limitation is not statistical. Because a human sets the priors, a prior can be set until a channel reports the return its owner was hoping for. A model built this way is a mirror: it can be angled to show the sponsor the result they wanted. This is a property of the method, not an accusation against any particular tool.
A last limitation is that the accuracy claims cannot be taken at face value, because every one is the maker's own. PyMC Labs published a 2025 benchmark showing its own tool beat Google's Meridian (PyMC Labs, 2025), one maker grading a rival; Google and Meta grade their own tools (Runge et al., 2024); and Recast advertises figures like better than ninety-five per cent accuracy with no outside audit (Recast, no date). It is the same conflicted-grader problem that appears whenever a platform marks its own homework (platforms grading their own ads; Gordon, Moakler and Zettelmeyer, 2023), and no standards body checks any of it. That no published challenge to these tools has yet appeared is a sign of the field's youth, not of agreement.
Open questions
The most important open question is whether the recommended cure actually works. The fix for weak identification is to calibrate the model against an experiment, yet no published study has closed the full loop and shown that a geographic experiment feeding into a prior recovers the true effect in the model's answer. The direction is sensible and widely recommended; the end-to-end proof does not yet exist, and part of it may live in the neighbouring work on geo experiments.
Underneath sits a broader unknown: with no neutral benchmark against a known ground truth, which of these tools is actually the most accurate is a question nobody can yet answer, and the peer-reviewed core that might settle it is thin, resting on one journal review (Migon et al., 2023) and a preprint systematic review (Lai, Cheng and Liu, 2026).
So what
One fact matters most in practice: in a weakly identified model, the analyst's assumptions, not the data, often decide the answer. That is what makes a marketing mix model useful, and also what makes it corruptible. A prior can be adjusted until the chart shows the channel someone wanted to look good. So the first thing to settle is who you are while running it. You are responsible for a number that a real budget will follow, not someone adjusting a tool until it shows what they want. The good news is that the rigorous move and the responsible move are the same move. A model adjusted to flatter is worthless as information, because it only tells you what you already believed, and the way to keep it honest is the same discipline that makes it accurate. Anchor the assumptions to a real experiment, and ask of every choice whether it serves the decision in front of you or the conclusion someone already preferred. When the campaign underperforms next quarter, a flattering model is the early-warning system you threw away.
For research practice
Treat the prior as the main result, not a preliminary step. Write your assumptions down before you see the result, keep a record of what you set and why, and report the uncertainty the model gives you rather than a single confident number. When two defensible sets of assumptions produce different answers, that gap is a result, not something to hide. This means the data alone cannot answer the question. The best thing you can do is compare against an experiment, so that at least one channel's effect is anchored to something real. A marketing mix model is one part of a broader measurement system, best used with other evidence rather than as the final answer, which we cover in triangulating measurement. And be realistic about the tools. Open-source code you can inspect is a real gain in transparency, but reading the code is not the same as having an independent party confirm the tool gives correct answers, and no such confirmation yet exists.
For companies
You will be told that an open-source, Bayesian marketing mix model tells you objectively and privately what your advertising did. Half of that is true. It is privacy-safe, and it is more inspectable than the black boxes it replaced. Question the word 'objective'. Because the assumptions can decide the answer, the same data can be made to support whichever channel someone inside the company favours. The risk is greatest when the person building the model also owns the budget it will justify. The safeguard is simple and works. Separate the people who build the model from the people whose spend it evaluates, insist on seeing the priors and the uncertainty rather than a single headline return, and calibrate against a real geographic test before you spend a large amount of money. When a vendor quotes a headline accuracy figure, ask who measured it. If the answer is the vendor, you have a marketing claim, not a verified result (Recast, no date). Turning these numbers into an actual decision is its own skill, covered in reconciling models into a decision.
For political parties
Campaigns spend across television, digital, mail and field, and every campaign wants to know which spending moved support. A marketing mix model is an appealing way to find out, and it does not require tracking individual voters, which is a genuine advantage in a privacy-conscious climate. The risk is specific to the setting. A consultant paid to run a channel can, with the same tool, set the assumptions that make that channel look decisive, and a campaign eager to justify its plan will not push back hard. That is not measurement. It is a rationale presented as measurement, and it fails after the money has been spent and the result is known. The rigorous version is the honest one here too. Fix the assumptions in advance, run a real geographic lift test where you hold spend in some regions and raise it in others, and treat the model as a way to read your own effort rather than to defend it. A number created to reassure the campaign manager is worthless to the candidate.
For government and policy
Governments are large advertisers, on public health, recruitment, safety and civic information, and they are accountable for whether that spending works. Marketing mix modelling offers a privacy-safe way to evaluate it, which fits the higher bar a public body should hold itself to. The same accountability also applies in the opposite direction. An agency that wants to show a campaign succeeded can commission a model whose assumptions deliver that verdict, and because no standards body certifies these tools, there is no external check to catch it. The safe approach is what a statistical office already does. Set the assumptions in advance, calibrate against a real experiment where the stakes justify it, publish the uncertainty rather than a single confident number, and keep a record of who built the model and what they assumed. For a regulator there is a second job. A headline accuracy claim in a measurement vendor's marketing is an assertion, not an audited fact, and treating it as fact is exactly the kind of claim consumer-protection rules exist to test.
How to use this
Before you trust a marketing mix model, ask four questions. Who set the priors or the regularisation, and do they benefit from the answer coming out a particular way? Were the assumptions written down before the result was seen, or adjusted until the chart looked right? Is any channel's effect anchored to a real experiment, or does the whole thing rest on historical spend that moved together? And who validated the tool, keeping in mind that the maker's own accuracy figure is not validation? A model that answers these well is a reliable way to understand what your advertising did. One that avoids these questions is a mirror, and you should trust it least when it tells you what you hoped to hear.
Case studies
Google's own tooling, from unofficial library to supported release. The clearest dated marker of the shift is Google's. Before 2022 the aggregate approach lived, at Google, in an unofficial open-source library the company did not formally support. That library was retired and archived in January 2026, its page redirecting users onward (Google, 2026). In its place Google released Meridian, reaching a first stable version early in 2025 as a supported, openly licensed Bayesian tool (Google, 2025). The migration is itself the evidence. The inspectable, supported form of this method did not exist before the privacy shift made aggregate measurement matter again. It is a documentation trail, not an independent test of whether the tool is accurate, and that distinction is the whole point of the article.
A named deployment, on Google's own product. In October 2025 Wharton's AI and Analytics Accelerator published a project, run with Google, that built a Meridian model on a few years of US sales data for Google's Pixel phones (Wharton AI and Analytics Accelerator, 2025). It is a rare dated, named use of one of these tools rather than a claim about them in the abstract, and it puts the article's central mechanism in plain view. The analysts set the model's flexibility and its prior beliefs about each channel's return by hand, and those choices moved the fit. The caveat is the one that runs through this piece. It is Google's tool, applied to Google's own product, in a partnership with Google, so it shows how the method is used rather than testing whether it is right.
Recast, and the vendor that grades itself. Recast is a commercial Bayesian marketing mix service that advertises high accuracy across a large number of client models (Recast, no date). It is a useful illustration of where the market has gone, and of the trap in it. The accuracy figures are self-reported, with no independent audit behind them, so they document what is claimed rather than what is confirmed. Used as a picture of where the commercial market has gone, Recast is informative. Used as evidence that the method is accurate, it is exactly the conflicted grader the field has yet to replace with a neutral one.
References
Dew, R., Padilla, N. and Shchetkina, A. (2024) 'Your MMM is broken: identification of nonlinear and time-varying effects in marketing mix models', arXiv preprint. Available at: https://doi.org/10.48550/arXiv.2408.07678 (Accessed: 18 August 2026).
Google (2025) Meridian [open-source software and documentation]. Available at: https://github.com/google/meridian (Accessed: 18 August 2026).
Google (2026) LightweightMMM [open-source software, deprecated]. Available at: https://github.com/google/lightweight_mmm (Accessed: 18 August 2026).
Gordon, B.R., Moakler, R. and Zettelmeyer, F. (2023) 'Close enough? A large-scale exploration of non-experimental approaches to advertising measurement', Marketing Science, 42(4), pp. 768–793. Available at: https://doi.org/10.1287/mksc.2022.1413 (Accessed: 18 August 2026).
Jin, Y., Wang, Y., Sun, Y., Chan, D. and Koehler, J. (2017) Bayesian methods for media mix modeling with carryover and shape effects. Mountain View, CA: Google Inc. Available at: https://research.google/pubs/bayesian-methods-for-media-mix-modeling-with-carryover-and-shape-effects/ (Accessed: 18 August 2026).
Lai, L., Cheng, Z. and Liu, Y. (2026) 'Trustworthy AI for marketing measurement: a systematic review of attribution, media mix modeling, and privacy-preserving methods', Research Square preprint. Available at: https://doi.org/10.21203/rs.3.rs-10322944/v1 (Accessed: 18 August 2026).
Marín, J. (2023) 'A new framework for marketing mix modeling: addressing channel influence bias and cross-channel effects', arXiv preprint. Available at: https://doi.org/10.48550/arXiv.2311.05587 (Accessed: 19 August 2026).
Meta (no date) Robyn [open-source software]. Available at: https://github.com/facebookexperimental/Robyn (Accessed: 18 August 2026).
Migon, H.S., Alves, M.B., Menezes, A.F.B. and Pinheiro, E.G. (2023) 'A review of Bayesian dynamic forecasting models: applications in marketing', Applied Stochastic Models in Business and Industry, 39(3), pp. 471–493. Available at: https://doi.org/10.1002/asmb.2756 (Accessed: 18 August 2026).
PyMC Labs (2025) PyMC-Marketing vs. Google Meridian [benchmark blog post], 8 September. Available at: https://www.pymc-labs.com/blog-posts/pymc-marketing-vs-google-meridian (Accessed: 19 August 2026).
PyMC Labs (no date) PyMC-Marketing [open-source software and documentation]. Available at: https://www.pymc-marketing.io/ (Accessed: 18 August 2026).
Recast (no date) Recast: Bayesian marketing mix modeling [product documentation]. Available at: https://www.getrecast.com/ (Accessed: 18 August 2026).
Runge, J., Skokan, I., Zhou, G. and Pauwels, K. (2024) 'Packaging up media mix modeling: an introduction to Robyn's open-source approach', arXiv preprint. Available at: https://doi.org/10.48550/arXiv.2403.14674 (Accessed: 18 August 2026).
Wharton AI and Analytics Accelerator (2025) Building marketing mix models with Google Pixel data using Meridian. Available at: https://ai-analytics.wharton.upenn.edu/student-programs/analytics-accelerator/building-marketing-mix-models-with-google-pixel-data-using-meridian/ (Accessed: 19 August 2026).
Wu, Y., Gu, Z., Deng, A., Zhu, J. and Chen, L. (2026) 'Hierarchical clustering as a novel solution to the notorious multicollinearity problem in observational causal inference', arXiv preprint. Available at: https://doi.org/10.48550/arXiv.2606.30992 (Accessed: 18 August 2026).
Explore the idea
Let’s talk
Invisible forces shape your world — until you hire Latenta®
Contact