Retail Media Measurement: The Store Sells the Ad and Grades It Too
Article M3-02
Retail media finally lets a retailer watch an ad turn into a sale, which is the closest thing market research has to a measurement lab. The catch is that the platform selling the ad also scores it, and the score it hands back systematically flatters the sale.
In brief
Retail media lets a retailer watch its own ad turn into an actual purchase, on its own checkout data. That closed loop is the best tool market research has for tying an ad to a sale: unlike a Google or social ad, which reports a modelled estimate of a conversion it cannot see, the retailer records the real receipt. The catch is that the party selling the ad also grades it, so its reported return on ad spend overstates the real effect, crediting sales that would have happened anyway. The mechanism is well evidenced; the retail-specific size of the overstatement is not. And this matters now for a reason unrelated to AI: as third-party tracking collapsed, the first-party loop became the last deterministic ad-to-sale measurement left standing, at new scale, forcing the standards that followed in 2024 and 2025. The fix never changes: trust the number only after a check the platform does not control.
What the theory says
The theory
Retail media is what you get when a shop that already knows what you bought starts selling ads next to the things you might buy next. Search a product on a large retailer's site or app, and some results are paid placements: the retailer runs an ad business on top of its store, and it is now very large. IAB puts commerce-media spending, the category it sits inside, past 150 billion dollars (IAB, 2025), an industry figure rather than an audited one, though the scale is not in doubt.
What makes retail media different is not tracking. Plenty of advertising is tracked: a Google or TikTok ad fires a pixel when someone clicks and later buys, and the platform reports a conversion. But that conversion is usually a modelled estimate, counted on the advertiser's own website, not in the shop. For the packaged-goods brands that dominate retail media, the real purchase happens on the retailer's shelf or app, where a search or social platform cannot see it at all. The retailer can.
The loop does not close the same way for every retailer. Joining the two moments, seeing the ad and buying the product, needs a single identity present at both. For a pure online store like Amazon, that identity is the account you are logged into when you see the ad and when you check out, so the loop closes by itself. Amazon has pushed this furthest: it shows ads not only in its store but across Prime Video, Twitch, Fire TV and, through its own ad network, on third-party sites and connected TVs, all tied to the same logged-in shopper and their purchases, the widest closed loop anyone has built. A supermarket like Tesco, Walmart or Kroger is the harder case, because most of its selling still happens in a physical shop, where you are anonymous at the till. The loyalty card supplies the missing identity: you scan it, and the itemised receipt is tied to you, the same shopper who saw the ad in the retailer's app or email. The card lets a bricks-and-mortar chain do what Amazon does natively. And nobody else holds both halves: a search or social platform saw the ad but not the supermarket till, and the brand sells through the shop but never sees the individual shopper. This connected view, from ad to real purchase, is the closed loop, and why people call retail media a measurement lab; the clicks and views it rests on are a research instrument in their own right, taken up in behavioural exhaust.
The obvious objection is that Google and Apple have their own ecosystems, so surely they close the same loop. They do close a loop, but a different one, and it helps to separate targeting, which ad to show you, from attribution, whether it worked. Neither uses your payment data to target ads; Apple gates that kind of tracking behind a consent prompt and makes refusing it a selling point. But both still measure whether their ads work, silently and inside their own walls. Apple attributes ad-driven app installs and clicks through SKAdNetwork and Private Click Measurement, and Google matches ad clicks to advertiser-reported conversions through Enhanced Conversions, using hashed data rather than card records. They close an ad-to-conversion loop without ever reading your wallet. What sets the retailer apart is not that it closes a loop but what its loop is made of: the itemised purchase itself, deterministic, on receipts the party doing the grading owns outright.
So the technique is not new; what changed is why it suddenly matters. Retail media networks existed before 2022 (Bartholomew and Williamson, 2022), and they use no measurement method that generative AI invented. Two post-2022 shifts made them decisive. The first is the collapse of third-party tracking: Apple's app-tracking prompt in 2021 and the long retreat from browser cookies broke the modelled attribution that search and social platforms rely on, pushing everyone else's ad-to-sale measurement towards estimates. The retailer's closed loop was immune, because it never needed third-party cookies; its identity is first-party, its own logged-in account or loyalty card. As the ordinary tools degraded, the retailer's became the last deterministic ad-to-sale measurement left standing. The second shift is scale. These networks grew from a curiosity into one of the largest advertising channels in the world in a few short years, so a measurement problem that used to be one company's quirk now shapes a large slice of all ad spending. As both shifts landed, the standards bodies acted. IAB and the Media Rating Council published their first retail media measurement guidelines in early 2024 (IAB and MRC, 2024). IAB then issued guidance on advanced measurement and data collaboration in October 2024 (IAB, 2024a) and on in-store measurement that December (IAB, 2024b), and dedicated guidance on incrementality in commerce media followed in November 2025 (IAB and IAB Europe, 2025). The dates matter, because they mark the moment the industry admitted its headline numbers needed rules.
The retailer that sells the ad is the same party that measures whether the ad worked, and the conflict here runs deeper than it does on a search or social platform. Google grades its own clicks, but it does not own the shop or sell a rival product. A retail media network does both. It sells the advertising, it owns the checkout that grades the result, and it often sells its own private-label brand on the same shelf as the advertiser it is grading. So the grade a campaign receives is written by the company that was paid to run it, on data only it holds, and sometimes by a direct competitor. No one has to cheat for this to cause a problem. It is a structural conflict: the party that can best see the result is also the party that benefits most from a favourable result.
Controversies
The current argument does not dispute that the loop measures. What it disputes is what retail media's favourite number actually means.
That number is usually a version of return on ad spend, built by attribution. Attribution credits a sale to an ad when the buyer saw the ad and then bought. This looks like proof, but it is not, because it cannot tell you what would have happened without the ad. A large share of the people who see a sponsored listing for a brand were going to buy that brand anyway. They searched for it. Crediting the ad with those sales counts them as new when they were not at risk of being lost. The real question is whether the ad caused the sale. Attribution answers an easier one: did a sale follow an ad.
The difference between those two questions has strong evidence, though not yet from inside retail media itself. Gordon, Moakler and Zettelmeyer (2023) compared observational attribution against randomised experiments at large scale and found that the attribution numbers systematically failed to recover the causal effect the experiments measured. That study looked at a large ad platform rather than a retailer's store, so the exact size of the overstatement for any named retail media network is not something a peer-reviewed study has measured. As far as this research could establish, no published paper puts a retailer's reported return on ad spend next to an incrementality benchmark and reports the difference. The mechanism is Tier A evidence, solid and general. The retail-specific size is unknown, and the honest position is to say so rather than make up a number.
Within retail media, the strongest peer-reviewed voice is Pauwels and Fagbola (2025), writing in the Journal of Retailing. They name the incrementality question directly, and they describe retailers having to balance their role as an advertising platform against their role as a competitor to the very brands advertising on them. That is the conflict-of-interest problem stated in an academic journal rather than in a rival's marketing.
And there is a rival's marketing, and it is loud. A whole category of independent measurement firms now sells itself on the claim that the platform's own return figure overstates the truth, and that only a proper experiment recovers the real effect. This is commercial positioning, not evidence, and it should be seen as one company arguing against another's number. It is worth knowing about because it shows the conflict is being argued in the open market, not only in theory. The same argument is starting to appear in fresh academic work. Li et al. (2026) propose an attribution method that corrects for cannibalisation, which is the effect of crediting an ad for sales it merely shifted rather than created. That paper is early, not peer reviewed, and has no citations yet, so it shows an active argument rather than a settled result.
Limitations
Suppose you do the honest thing and run an experiment. You hold out a group of shoppers from seeing the ad, run the ad to everyone else, and compare. Even then the measurement can mislead, in ways that are not obvious.
The first is divergent delivery. Modern ad platforms decide who sees an ad using an algorithm that targets the people most likely to buy. Braun and Schwartz (2025), writing in the Journal of Marketing, show that this optimisation can bias even a clean A/B test, because the treatment and control groups end up receiving the ad under different conditions the experiment never controlled. A well-run experiment can still produce a wrong result.
The second is that experiments in advertising are simply hard to run well. Johnson (2023) catalogues the ways field experiments in online advertising go wrong, from contaminated control groups to effects too small to detect against noisy sales. Experiments are not useless. The lesson is that the independent check comes at a price, and a badly run experiment can be more confident and more wrong than the attribution number it was meant to correct.
The third limitation is built into the closed loop itself. The loop only sees its own store. A shopper lives across many shops, many screens, and many weeks. An ad on one retailer can move a sale to a different retailer, or to next month, or from an online basket to a physical shelf. Lambrecht, Tucker and Zhang (2023) document exactly this kind of movement, showing how advertising shifts sales across channels and across time in ways a single seller's data cannot see. Older work on what happens when brands stop advertising points the same way (Phua et al., 2023). A retailer's closed loop only shows data from its own store, not the shopper's other activities.
There is also a reach problem hiding inside the clean story. The deterministic loop only holds where the retailer sees you logged in, on its own site or app, which for a supermarket is a minority of shoppers. Most buying is a card swipe in a physical aisle by someone who never opens the app or reads the email. To reach those shoppers the retailer goes offsite, matching its loyalty list onto Meta, Google or connected TV through a clean room, and there the tie between ad and purchase stops being a clean one-to-one link and becomes a probabilistic match whose accuracy depends on how many shoppers can be identified at all. The pristine closed loop is the narrow case; at the scale the budgets assume, much of the measurement is an identity match, not a logged-in fact.
The last limitation is about the word independent. The standards response leans on neutral measurement environments, often called clean rooms, where a brand's data and a platform's data can be compared without either side seeing the other's raw records. The machinery of those rooms is real and improving (Li et al., 2024), and it belongs to clean-room measurement rather than here. The catch is ownership. A clean room can be the platform's own environment. The largest retail media platform runs its measurement clean room on its own cloud. Neutral-sounding infrastructure does not make a neutral judge, and critical work on the industry's privacy turn reads the shift partly as self-serving rather than purely protective (McGuigan et al., 2023).
Open questions
Two things are genuinely unknown, and both matter to a buyer. The first is magnitude: nobody has published, for a named retail media network, how far its reported return on ad spend sits above what an incrementality experiment would credit. The direction is well supported; the size is not. Until someone runs that comparison, a buyer knows the grade is inflated but not by how much, which is the difference between distrusting a number and being able to correct it. The second is whether outside grading changes anything. Retail media platforms have begun submitting to Media Rating Council accreditation, and Instacart's advertising arm reported expanded accreditation across many retailer sites in late 2025 (Instacart, 2025); but no one has shown whether a platform's reported figures come down once an auditor is watching, and accreditation certifies process, not that the number fell. The fixes on offer share one shape, checking the platform's convenient number against something it does not control, whether a geo experiment comparing whole regions (Larson and Dotson, 2026) or a predict-then-confirm reconciliation from the team behind the original attribution finding (Gordon, Moakler and Zettelmeyer, 2026).
So what
Retail media really is the best setting market research has for tying an ad to a sale, on real purchases at a scale no panel reaches. The catch is that the party being graded produces the grade, so the number is not fully trustworthy. The rigorous move and the honest one are the same: treat the platform's figure as something to verify, not a final answer, and pay for at least one look from outside the loop.
For research practice
Treat the platform's return on ad spend as the vendor's self-assessment, because that is what it is. Your job is to recover the incremental effect, meaning the sales that would not have happened without the ad, and attribution does not give you that. The key step is to insist on a comparison the platform does not control: a holdout, a geo experiment, or a neutral reconciliation of the platform's number against an independent one. The details of that are in incrementality testing and the wider measurement stack. Two habits help you. First, ask what the reported number is counting, and specifically whether it counts buyers who were going to buy anyway. Second, remember the check's own limits, because a badly run experiment, biased by how the platform delivered the ad, can be as wrong as the attribution figure you distrusted. The aim is to hold every number, including your own, to the same question: does it hold up against a comparison the seller did not build.
For companies
You are the buyer, and you are being handed the seller's grade for the thing you bought. That grade decides where next quarter's budget goes, which makes it one of the most consequential numbers in the marketing function and one that is likely biased. The gain is real. Retail media can tell you which products sold after which placements, on actual receipts, which most media never could. The risk is that you scale spending on a figure designed to appear favourable. Before you reallocate budget on a platform's reported return, ask for the incremental version, and if the platform will not or cannot provide it, get one independent measurement from someone with no stake in the answer. Ask whether the measurement environment is genuinely third-party or the platform's own. Ask whether the number was checked by an accredited auditor, and, more sharply, whether the checked number differs from the unchecked one. A campaign that only looks good in the seller's own report has told you nothing you can bank.
For political parties
Campaigns buy commerce and retail media too, and they buy the same self-graded measurement with it. A vendor that reports strong reach or strong lift for your spending is reporting on its own performance, and the incentive runs exactly as it does for a commercial advertiser. The specific caution for a campaign is to keep the measurement question separate from the strategy question. A flattering placement report shows that a vendor delivered impressions and credited itself with whatever followed. Whether a message actually persuaded voters is a different question the report does not touch. If a supplier will not show you how a number was produced, or will not submit it to any check outside its own systems, treat that refusal as information. The discipline is the same one any serious buyer needs: the number that matters is the one that survives an independent look, not the one that arrives inside the invoice.
For government and policy
Two roles meet here. Public bodies buy advertising and its measurement, and they also set the rules that measurement runs under. As a buyer, a public agency spending on retail or commerce media should hold its suppliers to the incremental standard and to independent audit, for the same reason a company should, with the added duty that the money is public. As a rule-setter, the notable development is that the standards now exist at all. The industry's own bodies published measurement guidance and then incrementality guidance across 2024 and 2025, and an accreditation route now runs through the Media Rating Council. That is progress worth recognising, and also worth interrogating, because the standards are largely industry-written and accreditation certifies process rather than proving that reported effects are real. The useful question for policy is whether the graded number and the audited number agree, which is exactly the thing no one has yet published.
How to use this
Before you trust a retail media result, ask it four questions. Who produced this number, and do they get paid more if it looks good? Does it count sales the ad caused, or only sales that followed the ad? Has it been compared against something the platform does not control, such as a holdout or a geo experiment? And is the measurement environment genuinely independent, or the platform's own system with a neutral-sounding name? A number that passes all four is rare and worth acting on. A number that passes none is a vendor's opinion of its own work. Most retail media numbers sit somewhere in between, and knowing where a particular one sits is the main skill. The closed loop is a real instrument, the best that market research has for tying an ad to a sale. It is also controlled, for now, by the party with the most to gain, and the fix is to insist that someone outside the loop reads it too.
Case studies
Instacart's Carrot Ads and the Media Rating Council (2025). The clearest public example of a conflicted grader inviting an outside check is Instacart's advertising business, Carrot Ads, which reported expanded Media Rating Council accreditation across more than 240 retailer sites in late 2025 (Instacart, 2025). The accreditation is a genuine step, because it puts a platform's own advertising metrics in front of an independent auditor rather than leaving the grade entirely inside the seller. It is also a partial one, and reading it honestly means holding both halves. What the auditor accredited is worth reading closely: impressions, clicks, click-through rate, and viewable impressions. Those are counting metrics, the record of what was delivered and seen. The accreditation does not extend to return on ad spend, and it does not touch any claim about incremental sales. So the outside check lands squarely on the delivery numbers and stops short of the causal ones, which is the very place the conflict of interest bites hardest.
One Mighty Mill and the gap between reported and incremental (2026). A media agency published a controlled test for One Mighty Mill, a bread brand, that shows the shape of the problem inside a single account (QBR Media, 2026). One retail platform reported a return on ad spend of ten to fifteen times. The agency then moved the spend up and down in stages and watched what the actual sales did. Lifting spend by roughly 500 percent moved online units sold by only about 10 percent, which means the reported return was mostly crediting demand that was already there. A second platform in the same test showed genuine incremental response at a lower cost per unit, so the money was better spent there. Read this for its shape, not for a number you can carry to another network. It is an agency case study rather than a peer-reviewed result, the platforms are anonymised to protect the client, and one account is not a measured rate for any named network. What it illustrates is exactly the mechanism the rest of the article describes: subtract the sales that would have happened anyway, and a headline return can collapse.
References
Bartholomew, D.E. and Williamson, M. (2022) 'Retail media networks', Journal of Retailing and Consumer Services, 69, 103119. Available at: https://doi.org/10.1016/j.jretconser.2022.103119 (Accessed: 18 August 2026).
Braun, M. and Schwartz, E.M. (2025) 'Where A/B testing goes wrong: how divergent delivery affects what online experiments cannot (and can) tell you', Journal of Marketing, 89(2), pp. 71–95. Available at: https://doi.org/10.1177/00222429241275886 (Accessed: 18 August 2026).
Gordon, B.R., Moakler, R. and Zettelmeyer, F. (2023) 'Close enough? A large-scale exploration of non-experimental approaches to advertising measurement', Marketing Science, 42(4), pp. 768–793. Available at: https://doi.org/10.1287/mksc.2022.1413 (Accessed: 18 August 2026).
Gordon, B.R., Moakler, R. and Zettelmeyer, F. (2026) Predicted incrementality by experimentation (PIE) for ad measurement. NBER Working Paper 35044. Cambridge, MA: National Bureau of Economic Research. Available at: https://doi.org/10.3386/w35044 (Accessed: 18 August 2026).
IAB (2024a) Retail media advanced measurement and data collaboration guidelines. 31 October. New York: Interactive Advertising Bureau. Available at: https://www.iab.com/guidelines/retail-media-advanced-measurement-and-data-collaboration/ (Accessed: 19 August 2026).
IAB (2024b) Retail media in-store measurement guidelines. 3 December. New York: Interactive Advertising Bureau. Available at: https://www.iab.com/guidelines/in-store-retail-media/ (Accessed: 19 August 2026).
IAB (2025) Demystifying incrementality in commerce media. 9 September. New York: Interactive Advertising Bureau. Available at: https://www.iab.com/guidelines/demystifying-incrementality-in-commerce-media/ (Accessed: 19 August 2026).
IAB and IAB Europe (2025) Guidelines for incremental measurement in commerce media. 3 November. New York and Brussels: Interactive Advertising Bureau and IAB Europe. Available at: https://www.iab.com/guidelines/guidelines-for-incremental-measurement-in-commerce-media/ (Accessed: 19 August 2026).
IAB and MRC (2024) Retail media measurement guidelines. Final version, 29 January. New York: Interactive Advertising Bureau and Media Rating Council. Available at: https://www.iab.com/insights/retail-media-measurement-guidelines/ (Accessed: 19 August 2026).
Instacart (2025) Instacart receives expanded MRC accreditation for Carrot Ads, bringing verified ad metrics to more than 240 ecommerce sites. Press release, 6 November. Available at: https://www.prnewswire.com/news-releases/instacart-receives-expanded-mrc-accreditation-for-carrot-ads-bringing-verified-ad-metrics-to-more-than-240-ecommerce-sites-302606493.html (Accessed: 19 August 2026).
Johnson, G.A. (2023) 'Inferno: a guide to field experiments in online display advertising', Journal of Economics & Management Strategy, 32(3), pp. 469–490. Available at: https://doi.org/10.1111/jems.12513 (Accessed: 18 August 2026).
Lambrecht, A., Tucker, C. and Zhang, X. (2023) 'TV advertising and online sales: a case study of intertemporal substitution effects for an online retailer', Journal of Marketing Research, 61(2), pp. 248–270. Available at: https://doi.org/10.1177/00222437231180171 (Accessed: 18 August 2026).
Larson, J.S. and Dotson, J.P. (2026) 'Prospecting versus retargeting: insights from a geography-based field experiment', Journal of Interactive Marketing, 61(3), pp. 257–274. Available at: https://doi.org/10.1177/10949968261422740 (Accessed: 18 August 2026).
Li, D., Yuan, B., Yang, Z., Chen, Q. and Song, L. (2026) Attributed, but not incremental: cannibalization-corrected attribution for large-scale advertising. arXiv:2606.26690. Available at: https://doi.org/10.48550/arXiv.2606.26690 (Accessed: 18 August 2026).
Li, K., Chen, X., Leng, L., Xu, J., Sun, J. and Rezaei, B. (2024) 'Privacy preserving conversion modeling in data clean room', in Proceedings of the 18th ACM Conference on Recommender Systems (RecSys '24), pp. 819–822. Available at: https://doi.org/10.1145/3640457.3688054 (Accessed: 18 August 2026).
McGuigan, L., West, S.M., Sivan-Sevilla, I. and Parham, P. (2023) 'The after party: cynical resignation in Adtech's pivot to privacy', Big Data & Society, 10(2). Available at: https://doi.org/10.1177/20539517231203665 (Accessed: 18 August 2026).
Pauwels, K. and Fagbola, L. (2025) 'Understanding retail media: perspectives and implications for stakeholders', Journal of Retailing, 101(3), pp. 315–330. Available at: https://doi.org/10.1016/j.jretai.2025.08.005 (Accessed: 18 August 2026).
Phua, P., Hartnett, N., Beal, V., Trinh, G. and Kennedy, R. (2023) 'When brands go dark: a replication and extension: examining market share of brands that stop advertising for a year or longer', Journal of Advertising Research, 63(2), pp. 172–184. Available at: https://doi.org/10.2501/jar-2023-009 (Accessed: 18 August 2026).
QBR Media (2026) How we used incrementality testing to optimise One Mighty Mill's retail media budget. Case study. Available at: https://qbrmedia.com/one-mighty-mill-retail-media-incrementality-testing (Accessed: 19 August 2026).
Explore the idea
Let’s talk
Invisible forces shape your world — until you hire Latenta®
Contact