Data Provenance: A Number Can Now Come From Nowhere
Article M0-05
When a machine helps produce a finding, some claims trace back to real data and some trace back to nothing. Telling them apart is the job that used to be automatic.
In brief
For most of research history you could assume a number was measured. Somebody asked people, counted something, ran the analysis, and even a bad number came from a real measurement you could walk back to. A generative model breaks that assumption. It can produce a fluent, specific, correctly formatted number, quote or citation that corresponds to nothing at all. Two peer-reviewed studies found that when ChatGPT was asked to support its claims, roughly half of the older model's citations were entirely fabricated, and the response the field is building, forcing every machine-written claim to cite its source, only half works: an audit of four AI search engines found that a quarter of their citations did not support the sentence they were attached to. Provenance is the discipline of tracing each claim back through the chain that produced it, and it has moved from bookkeeping to the main quality control.
What the theory says
The theory
Everyone who reads a research finding makes a quiet assumption: the number was measured. Somebody asked real people, or counted real events, or ran a real model on real inputs, and the number on the slide stands for something that happened. The assumption is so basic that nobody states it. It is the reason a finding means anything.
Provenance is the record behind that assumption. It is the chain from the thing that happened in the world to the number in the deck: which data, collected how, from which people, put through which analysis, written up by whom. For most of the field's history provenance was dull and mostly implicit. You did not usually walk the chain, because you did not have to. A number came from a measurement, even a sloppy or biased one, so in principle you could always trace it back to something real.
A generative model breaks the assumption at its root. It can produce a number, a statistic, a quotation or a citation that was never measured, never said, and never published, and it produces it in the same confident, well-formatted prose as a real one. This is what the research literature calls hallucination. The standard survey of the problem defines it as a model generating "unintended text" that reads fluently while failing to correspond to any real source (Ji et al., 2023).
The most useful evidence that this is real, and not a rare glitch, comes from studies that simply asked a model for its sources and checked them. Walters and Wilder (2023), in Scientific Reports, had ChatGPT write 84 short literature reviews and examined all 636 references it produced. They report that "55% of the GPT-3.5 citations but just 18% of the GPT-4 citations are fabricated", meaning the cited paper does not exist, and that many of the real citations still carried substantive errors. A medical study found the same pattern harder: across 115 references in generated medical content, Bhattacharyya et al. (2023) report that "47% were fabricated, 46% were authentic but inaccurate, and only 7% were authentic and accurate." Both are peer-reviewed, and both are measuring the same new thing: a source that was invented rather than mistaken.
That distinction is the whole point, and it is worth being precise about. The classic way of accounting for error in research, the total-survey-error tradition, sorts every mistake into a family: coverage error from who you could reach, sampling error from who you happened to draw, non-response error from who declined, measurement error from a question that misfired (Groves and Lyberg, 2010). Every one of those is a distortion of a real value. A hallucination distorts nothing, because there was no real value to begin with. A fabricated citation is not a bad measurement. It is the absence of one. The old taxonomy has no cell for it, which is exactly why provenance stops being clerical and becomes the check that matters: you can no longer assume the number came from a measurement, so you have to be able to show that it did.
The field's answer has a name too. It is citation-grounded reporting, and the principle is simple: make every machine-produced claim point at a specific source, and check that the source actually supports it. The formal version is the AIS framework, short for Attributable to Identified Sources, which stipulates that any generated statement about the world should be "verified against an independent, provided source" (Rashkin et al., 2023). The engineering version is retrieval-augmented generation, where the model is made to fetch real documents first and answer only from them, cite them, and leave a trail. Its own designers describe it as the response to a model's "non-transparent, untraceable reasoning" (Gao et al., 2023). There are now automated tools that score how well a grounded answer sticks to its sources (Es et al., 2024). The shape of the fix is agreed. Whether it works is the argument.
Controversies
The live question is whether citation-grounded reporting delivers what it promises, and the honest answer is that it half does.
The sharpest evidence is an audit that took the promise at face value and measured it. Liu, Zhang and Liang (2023) had humans check four popular AI search engines, the kind that answer a question and show citations beside each sentence. They found that "a mere 51.5% of generated sentences are fully supported by citations and only 74.5% of citations support their associated sentence", despite what they call the systems' "facade of trustworthiness". Read the second number slowly, because it is the one that matters here. A system built specifically to cite its sources still attaches a quarter of its citations to sentences those sources do not support. The citation is present. It is just wrong.
A larger and more recent audit found worse. The Tow Center for Digital Journalism tested eight AI search tools and reported that collectively they answered more than 60% of source-attribution queries incorrectly, naming the wrong outlet or a source that did not carry the claim, and that the paid premium versions did no better than the free ones (Jaźwińska and Chandrasekar, 2025). This is trade-press research rather than a peer-reviewed study, so treat the exact figure as indicative, but it points the same way as Liu and colleagues. A citation sitting next to a claim is not proof that the claim came from there. The link has to be walked, not trusted.
Around that sits a pattern that keeps showing up wherever AI reaches research: the market sells the solution faster than anyone has shown it works. Research platforms now advertise "insights 100% grounded in human truth" and "citation-backed insights", and those are claims, not evidence (GWI, 2026). The most useful case is one where a vendor published a number and someone independent checked it. The research assistant Elicit reports roughly 96% accuracy and a 1% hallucination rate and grounds each claim to a specific sentence in a source. The one independent peer-reviewed evaluation of it measured 81.4% extraction accuracy against a human benchmark of 86.7% (Hilkenmeier et al., 2025). The tool is genuinely useful and the gap between the advertised number and the audited number is the entire subject of this article. How to run that check yourself is the vendor-audit procedure covered in auditing a synthetic vendor, and the general discipline of scoring a tool before trusting it is covered in benchmarks and evals.
Limitations
The evidence here is strong on the problem and thinner on the fix, and where it is thin the reader should know it.
The fabrication rates date fast. The most cited figure in this piece, 55% falling to 18% between two model versions, is itself a moving target, and newer models fabricate less. That is real progress and it is also the trap: a number that improves every release invites you to wait for it to reach zero. It has not reached zero, and the class of error survives even as the rate falls, which is why the answer is to trace claims rather than to trust the latest model. The rate tells you how often you will need the provenance. It does not remove the need for it.
Almost all of the hard measurement comes from next door, not from market research. The fabrication studies are in medicine, academic literature search and law, because that is where people have sat down and checked every reference. As far as this research could establish, nobody has published a fabrication rate for numbers inside a commissioned survey-research deliverable. The mechanism plainly transfers, since the same models write the same way whatever the topic, but the specific rate for our field is not yet on record, and the honest move is to say so rather than borrow a medical number and present it as ours.
Even the checking tools have a blind spot worth naming. Many automated grounding scores are themselves produced by a language model judging whether one passage supports another, so the auditor can share the failure modes of the thing it audits. Grounding is being instrumented, which is not the same as being solved.
Open questions
No one has measured how often a hallucinated number reaches a finished research deliverable. As far as this research could establish, the rate is unpublished for commissioned survey research specifically. The nearest public evidence is from consulting and assurance work, where a fabricated citation surviving into a signed report is now documented (see the case studies), but the survey-research version of that number does not exist yet.
Whether provenance standards will reach the research pipeline is unknown. The infrastructure exists in adjacent worlds. Content Credentials, a cryptographic standard that records where a media file came from and whether AI touched it, is maturing (C2PA, 2026), and there is a mature, generic vocabulary for expressing the lineage of any claim as data (W3C, 2013). Neither was built for the chain that runs from a survey to a slide, and there is no published evidence yet that research vendors are adopting either. Blockchain provenance for research data has been proposed for years with no real uptake to show for it, for the same reason: the idea is sound and the adoption is not there.
Whether grounding closes the gap or just moves it is unsettled. The audits above measured systems that already ground their answers and still leave a residue of unsupported claims. Nobody has shown that the residue trends to zero as the tools improve, rather than shrinking to a smaller, harder-to-spot fraction. A smaller error rate that everyone trusts more can be more dangerous than a larger one nobody believes.
So what
The constraint has moved, the same way it moved for sample quality. When every number came from a measurement, provenance was implicit and you rarely had to check it. Now that a machine can produce a number with nothing behind it, provenance has to be explicit, and it has to be checked rather than assumed. Everything below follows from that one shift.
For research practice
Start with the uncomfortable symmetry. This article is a description of how a fabricated claim slips into a real deliverable, and a description of that is also a recipe for doing it on purpose. An undisclosed synthetic top-up in a thin subgroup, or an AI-written summary that rounds a soft finding into a hard one, is a way to launder a number that was never measured into a report that looks fully sourced. Naming the move is the only way a reader learns to catch it, and it puts the obligation on whoever holds the number.
Your job is to be able to show where each number came from, not to produce a deck that reads well. Treat every machine-touched claim as unverified until you can walk it back to a source you can open. And open it, because a citation is not a source: a quarter of the citations in systems built to cite are wrong (Liu, Zhang and Liang, 2023). Keep the lineage as you work. Which data, which model, which prompt, which analysis produced this figure. When one of those hops is a generative step, that is the first place to look when a number seems too clean. Where part of the input was synthetic, say so in the deliverable rather than smoothing it in, because why a synthetic number can be wrong is its own subject, covered in validity for synthetic evidence, and the duty to disclose it is covered in disclosure.
For companies
The practical question is what to require from a supplier who used AI somewhere in the work, and it is a short list.
Ask for the provenance chain. For a headline number, which parts came from data and which from a model, and can they show you the source behind it. Then open the source and check that it says what the claim says. Treat "our AI is grounded in real consumer data" as a marketing line rather than an answer (GWI, 2026), and ask for the audited grounding rate instead of the advertised one, because the two can differ by fifteen points on the same tool (Hilkenmeier et al., 2025). Scoring a tool honestly before you rely on it is covered in benchmarks and evals.
There is a cost worth naming. Checking provenance takes time that used to go into producing more analysis, and a smaller finding you can trace is worth more than a bigger one you cannot. A claim you cannot source is a different kind of object from one you can, not a cheaper version of it, and it should carry a different level of confidence in the room. Whether a number you can trace is then worth acting on is the buyer's decision rule, covered in decision-grade evidence.
For political parties
This is where the stakes are highest and the incentives are worst, so the ethical line comes first.
A statistic that agrees with your existing strategy is now the cheapest possible thing for someone to manufacture, and internal research is exactly where nobody checks a number that says what the room wants to hear. The methodological argument and the ethical one point the same way here, which is the strongest thing this subject has going for it. A finding you cannot trace to a real measurement is intelligence you cannot use, because you have no way to tell it apart from a fabrication that flatters you. Provenance is the early-warning system you are paying for, and a number with no source destroys it.
Defensively, the same habit lets you read someone else's evidence. When a convenient poll or a striking "study" appears, ask who produced the underlying data, what the claim is sourced to, and whether that source says what the headline says. Those questions are answerable, and reluctance to answer them is itself informative.
For government and policy
Public bodies commission and publish research, and both the buying and the publishing now need provenance built in. The risk is no longer hypothetical. Fabricated citations have reached signed, paid deliverables, including a government assurance review that quoted a court judgment that did not exist (see the case studies). A report whose credibility rests on its citations is now resting on something a machine can invent.
Procurement is the strongest lever, and it is available without waiting for a standards body. Specify that any AI-assisted deliverable carries a traceable source for every material claim, and spot-check a sample of them on delivery rather than taking the citations on trust. The rulebook is beginning to catch up: the 2025 ICC/ESOMAR International Code now requires disclosure when AI or synthetic data was used in producing a finding, covered in the 2025 Code, and the industry's buyer checklist already frames provenance and traceability as questions to put to a supplier (ESOMAR, 2023). The buyer who asks them is doing the thing the standard is still learning to require.
How to use this
Four questions, in order, to trace any claim a machine helped produce. They apply to a supplier's report, a colleague's slide, or a number in the news.
- Can you show me the source behind this number? If it came from an AI summary, what did the summary summarise? No answer means the chain is already broken.
- Open the source. Does it actually say this? A citation is a pointer, and in grounded systems the pointer is wrong about a quarter of the time. Reading it is the check.
- Which steps were generative? Where a model wrote, summarised, or filled a gap, a number can enter that came from nothing. Those hops get the scrutiny.
- Was any of it synthetic, and was that disclosed? An undisclosed synthetic step is an untraceable hop, and an untraceable hop is where a fabricated number hides.
A claim that survives all four is one you can stand behind. A claim that breaks at the first is directional at best, and often less. The mistake is not using a number you cannot fully trace, which is sometimes unavoidable. The mistake is not knowing which kind you are holding.
Case studies
Mata v. Avianca (2023). A lawyer preparing a federal filing asked ChatGPT for supporting case law and submitted what it produced: six fabricated judicial opinions, complete with invented quotations and fake internal citations. The court sanctioned the attorneys $5,000 and found bad faith (Judge Castel, sanctions order 22 June 2023). It is the clearest set-piece of a hallucinated citation surviving all the way into a signed, consequential deliverable, and it is why the legal profession learned this lesson first. The count of fabricated cases here follows the legal reporting on the matter rather than the court's docket directly.
Deloitte's DEWR assurance review (2025). This is the nearest thing yet to the market-research version. Deloitte Australia delivered a paid independent assurance review to a government department that contained a fabricated quote from a Federal Court judgment and references to academic papers that do not exist, produced with a generative model. The firm refunded part of the roughly A$440,000 fee after the errors were found. It is a commissioned analytic deliverable, the kind an insights team ships, with a hallucinated citation traced back to nothing and a dollar cost attached (OECD.AI, 2025).
The Tow Center AI-search audit (2025). Researchers at Columbia's journalism school tested eight AI search tools on their ability to attribute a claim to the right source and found them wrong more than 60% of the time, with the paid versions no better than the free (Jaźwińska and Chandrasekar, 2025). It is the cleanest demonstration that an answer with citations is not the same as an answer with sources, which is the practical heart of provenance.
References
Bhattacharyya, M., Miller, V.M., Bhattacharyya, D. and Miller, L.E. (2023) 'High rates of fabricated and inaccurate references in ChatGPT-generated medical content', Cureus, 15(5), e39238. Available at: https://doi.org/10.7759/cureus.39238 (Accessed: 11 August 2026).
C2PA (2026) Coalition for Content Provenance and Authenticity: technical specification (version 2.x). Available at: https://spec.c2pa.org (Accessed: 11 August 2026).
Es, S., James, J., Espinosa-Anke, L. and Schockaert, S. (2024) 'RAGAS: automated evaluation of retrieval augmented generation', Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, pp. 150–158. Available at: https://doi.org/10.18653/v1/2024.eacl-demo.16 (Accessed: 11 August 2026).
ESOMAR (2023) 20 questions to help buyers of AI-based services for market research and insights. Available at: https://esomar.org/publications/esomar-20-questions-to-help-buyers-of-ai-based-services (Accessed: 11 August 2026).
Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., Wang, M. and Wang, H. (2023) 'Retrieval-augmented generation for large language models: a survey', arXiv preprint arXiv:2312.10997. Available at: https://doi.org/10.48550/arXiv.2312.10997 (Accessed: 11 August 2026).
Groves, R.M. and Lyberg, L. (2010) 'Total survey error: past, present, and future', Public Opinion Quarterly, 74(5), pp. 849–879. Available at: https://doi.org/10.1093/poq/nfq065 (Accessed: 11 August 2026).
GWI (2026) Spark: ground your AI in human truth. Available at: https://www.gwi.com/platform/spark (Accessed: 11 August 2026).
Hilkenmeier, F., Pelzer, M., Stierle, C. and Fink-Lamotte, J. (2025) 'Evaluating the AI tool "Elicit" as a semi-automated second reviewer for data extraction', Social Science Computer Review. Available at: https://doi.org/10.1177/08944393251404052 (Accessed: 11 August 2026).
International Chamber of Commerce and ESOMAR (2025) ICC/ESOMAR international code on market, opinion and social research and data analytics. 5th edn. Released 29 September. Available at: https://standards.esomar.org/resources/international-codes (Accessed: 11 August 2026).
Jaźwińska, K. and Chandrasekar, A. (2025) 'AI search has a citation problem', Columbia Journalism Review (Tow Center for Digital Journalism), 6 March. Available at: https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php (Accessed: 11 August 2026).
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y.J., Madotto, A. and Fung, P. (2023) 'Survey of hallucination in natural language generation', ACM Computing Surveys, 55(12), pp. 1–38. Available at: https://doi.org/10.1145/3571730 (Accessed: 11 August 2026).
Liu, N.F., Zhang, T. and Liang, P. (2023) 'Evaluating verifiability in generative search engines', Findings of the Association for Computational Linguistics: EMNLP 2023, pp. 7001–7025. Available at: https://doi.org/10.18653/v1/2023.findings-emnlp.467 (Accessed: 11 August 2026).
Mata v. Avianca, Inc. (2023) 678 F. Supp. 3d 443 (S.D.N.Y.). Available at: https://en.wikipedia.org/wiki/Mata_v._Avianca,_Inc. (Accessed: 11 August 2026).
OECD.AI (2025) AI Incidents Monitor: Deloitte Australia report for DEWR contained AI-generated fabricated references and a fabricated court quote. Incident 2025-10-05-be45. Available at: https://oecd.ai/en/incidents/2025-10-05-be45 (Accessed: 11 August 2026).
Rashkin, H., Nikolaev, V., Lamm, M., Aroyo, L., Collins, M., Das, D., Petrov, S., Tomar, G.S., Turc, I. and Reitter, D. (2023) 'Measuring attribution in natural language generation models', Computational Linguistics, 49(4), pp. 777–840. Available at: https://doi.org/10.1162/coli_a_00486 (Accessed: 11 August 2026).
Walters, W.H. and Wilder, E.I. (2023) 'Fabrication and errors in the bibliographic citations generated by ChatGPT', Scientific Reports, 13, 14045. Available at: https://doi.org/10.1038/s41598-023-41032-5 (Accessed: 11 August 2026).
W3C (2013) PROV-DM: the PROV data model. W3C Recommendation, 30 April. Available at: https://www.w3.org/TR/prov-dm/ (Accessed: 11 August 2026).
Explore the idea
Let’s talk
Invisible forces shape your world — until you hire Latenta®
Contact