Research Archives and Retrieval: The Archive That Answers Back

Article M5-01

Generative models now read across an entire research archive and write a single cited answer, a genuinely new thing that enterprise search never produced. That answer can also be fluent, confident and wrong, and no vendor of an insight archive publishes a rate for how often that happens.

In brief

For twenty years, searching a company's research archive meant getting back a list of documents and reading them yourself. Since 2023, a handful of market-research vendors have shipped something different. You ask a question in plain language, and a generative model reads across every report, deck and transcript it can retrieve, then writes you one synthesised answer with footnotes pointing back to the sources. That answer is a new artefact, and it carries a new failure. It can be fluent and confident and still not match the documents it cited, or cite nothing at all. Everyone building these tools agrees the fix is a visible citation the reader can check. The awkward part is the evidence. The one independent benchmark of enterprise archive search finds even the best systems score low, mostly because they answer from evidence they never actually found. And no vendor of a commercial insight archive publishes an independent rate for how faithful its answers are, which is the single number a buyer would most want to see.

What the theory says

The theory

Ask a research team a question they have answered before, and the honest reply is often that the answer is in a folder somewhere. A company that has commissioned research for a decade owns thousands of reports, decks, transcripts and tracker waves. Until recently, searching that pile meant the same thing it meant in 2005: you typed some words, the system returned a list of documents, and you opened them and read. Enterprise search found the file. It did not answer the question.

What changed after 2022 is that a generative model can now read across the whole archive and write the answer itself. You ask, in plain language, what do our studies say about why lapsed subscribers leave?, and the system retrieves the relevant passages from wherever they sit and composes a single synthesised answer, with footnotes pointing back to the specific reports it drew on. The underlying machinery is retrieval-augmented generation, where a model is fed retrieved source passages and writes grounded on them rather than on its training memory (Gao et al., 2023).

What matters here is the output. The synthesised answer is a genuinely new artefact. It is not the document enterprise search returned, and it is not a chatbot's free-floating opinion. It is a written answer to your question, assembled on demand from your own evidence.

This is shipping, from named vendors, and has been since 2023. Market Logic launched DeepSights that year as an assistant over a market-intelligence archive (Market Logic Software, 2023). Stravito announced its Assistant in April 2024, offering natural-language questions across an owned archive with page-level references and a framing it calls Glass Box AI (Stravito, 2024). Dovetail shipped Ask Dovetail in October 2024, built on Claude 3.5 Sonnet, with what it describes as Wikipedia-style citations back to the specific data point (Dovetail, 2024). The exact capability this article is about is on sale from more than one company. Even withing common tools like Google Workspaces.

The reason to take it seriously rather than treat it as a demo is that grounding a model's answer in retrieved internal sources genuinely does improve it. Karakurt and Akbulut (2025) reviewed 63 primary studies of retrieval-augmented generation for enterprise knowledge work and found consistent gains in accuracy and relevance over a model answering from memory alone. That is a peer-reviewed systematic review, the strongest single piece of evidence in this article, and it says that the basic idea is sound.

The catch is that it can also fail in a new way. A synthesised answer can be fluent, confident and wrong. It can state something the retrieved documents do not support, blend two studies into a claim neither of them made, or cite a source that does not actually back the sentence it is attached to.

This is why everyone building these tools has converged on the same safeguard: the citation. An answer you can trace back to the report it came from is one you can check, and an answer with no visible source is one you have to take on faith. General enterprise vendors make the same point, that an AI answer without a citation cannot be verified (Glean, 2024).

Whether those citations actually hold up, whether the cited source really supports the sentence, is itself measurable, and whether an answer's citations support it is a question researchers have learned to score. Liu, Zhang and Liang (2023), testing generative search engines, found that only about half of the sentences these systems produced were fully supported by their own cited sources, and only about three-quarters of citations actually backed the claim they were attached to. Tools now exist to measure this automatically (Es et al., 2024). The check is real. The question the rest of this article asks is whether anyone selling you an archive assistant has run it where you can see.

Controversies

The debate is not about whether the capability exists. It is whether the deployed version works well enough to base a decision on.

Choubey et al. (2025) provide the strongest independent evidence. They built a benchmark for deep search over heterogeneous enterprise data: 39,190 real enterprise artifacts of mixed type, and a set of questions that require finding and combining evidence scattered across them. This is the closest public proxy for the archive-QA task an insight team actually faces. Even the best agentic systems, the ones that plan multi-step searches rather than retrieving once, reached an average score of only around 33 out of 100. The authors traced most of the failures to retrieval. The evidence needed to answer was not found, so the model was left reasoning over an incomplete picture and filling the gap. That is the exact mechanism behind a confident wrong answer. The model does not write worse when it has half the evidence than when it has all of it. It writes just as fluently, which is the problem.

Bruckhaus (2024) makes the same point more bluntly in a paper whose title states that retrieval-augmented generation does not work for enterprises. The argument is that the way these systems are usually built runs into security, accuracy, scale and integration problems that demos never surface. The author is affiliated with a vendor in the space, so this is a competitor's critique rather than a neutral audit, and it is worth reading as constructive rather than disinterested. Together, the two make a consistent point. The gap between a working demo and a dependable deployment is wide, and it is mostly about whether the right evidence gets retrieved in the first place.

Every claim a vendor makes about its own grounding is self-reported. DeepSights is described as checking all of the available sources for an answer (Market Logic Software, 2023). Dovetail's marketing has pointed to large weekly time savings for its users (Dovetail, 2024). These may well be true. But a design assertion that a system checks its sources is not a measured rate of how often it gets them right, and a headline productivity figure with no study behind it is marketing, not evidence. Where a number cannot be traced to the study or dataset that produced it, this article treats it as a claim about what the vendor intends, not a fact about what the tool delivers. The distinction matters most precisely when the number is impressive.

Limitations

Two limits shape how far this can be trusted today, and a third limits what the systems can even read.

The first is that most of these deployments are younger and rougher than the marketing suggests. Brehme et al. (2025) interviewed 13 practitioners building retrieval-augmented systems in industry and found that most deployments were still prototype-stage question-answering for a specific domain, and that evaluation was overwhelmingly manual. Teams were checking outputs by hand, reading answers and judging them, because automated faithfulness scoring was not yet part of how they worked. That has a direct consequence for a research team. The check that makes a synthesised answer safe, verifying that each claim traces to a source that supports it, is real but is not yet running automatically at the scale an archive assistant operates. Somebody still has to read.

The second is a failure mode specific to a research archive, as opposed to a generic document store. An archive accumulates over years, which means it holds studies that have since been superseded, findings that were later contradicted, and two trackers that disagree. A model retrieving across all of it can surface a stale result as though it were current, or blend two conflicting numbers into a false consensus. Vendors are aware of this. Stravito, for one, says its Assistant flags when older source material is used and points out conflicting data points (Stravito, 2024). That the vendors now name the problem is a good sign. It is not, on its own, evidence the flagging works. No independent test of the flagging has been published.

The third limit is simpler and easy to forget. A great deal of market research lives in formats a model reads poorly: long PDF reports, slide decks where the argument is in the layout, tables, and charts whose message is never written out in words. An assistant that retrieves and reasons over clean text can miss the finding that only exists as a bar in a chart on slide 34. Reading these formats reliably is an active research problem, not a solved one, so an archive that is mostly decks and PDFs is exactly the archive these tools handle least well.

Open questions

Will any vendor of a commercial insight archive publish an independent faithfulness or citation-support rate for its own product? None does today. Researchers have measured how often a research-grade generative search engine cites its sources correctly, in the open (Liu, Zhang and Liang, 2023). For DeepSights, Stravito Assistant or Ask Dovetail, none of them reports such a rate, and no third party has measured it. The number a buyer would most want, how often the answer faithfully reflects the documents it cites, is the one number the market does not provide.

What standard would a buyer use to demand that number? None exists yet. Insights Association and ESOMAR have general guidance on using AI in research, but neither publishes a benchmark or certification for archive-answer grounding. A buyer comparing two tools therefore has no shared measure to ask each vendor to report. The independent evidence that does exist stops at the generic layer, one step short of any specific product.

If no product-level rate exists, how close does the best public proxy get? The closest public, dated, traceable faithfulness figure comes from the Vectara hallucination leaderboard, which as of its 11 May 2026 update put the best models at roughly 2 to 3 percent hallucination (Vectara, 2026). Two caveats travel with that number. It measures a model summarising a single document, not synthesising across a whole archive, which is a much harder task. And it is published by a company that sells retrieval tooling, so it is not a disinterested audit. It is the best public proxy available and still only a proxy. The honest state of the field is that the capability ships, the check that would make it trustworthy is well understood, and the rate at which any given product passes that check is not something you can currently look up.

So what

Your archive can now answer questions instead of just returning documents, and that is a real gain that more than one vendor can sell you today. The answer is a new artefact, assembled on demand, and it comes with a new way to be wrong: fluent, confident and not actually supported by the sources it names. The safeguard everyone agrees on is the visible citation, which turns a claim you would have to trust into one you can check. No vendor gives you a number for how often the citation holds, so the burden of checking sits with you. The same burden follows any AI-written deliverable, including a study an agent runs from question to finished deck on its own (letting an agent run the study).

For research practice

Treat the assistant as a fast first draft of an answer, never as the answer. Its real value is speed and recall: it reads across more of the archive than you would have time to, and it surfaces the reports you had forgotten you owned. Use it for that. Then do the thing the tool cannot yet reliably do for itself: open the citations and confirm that each load-bearing claim actually sits in the source it points to. The evidence says this matters. When generative search engines were tested in the open, a meaningful share of their sentences were not fully supported by their own citations (Liu, Zhang and Liang, 2023). No commercial archive tool has shown it does better. The failure to guard against is the confident summary of a question your archive cannot actually answer. Choubey et al. (2025) found that the most common way these systems fail is by reasoning over evidence they never retrieved, which reads, on the page, exactly like a well-sourced finding. A quick way to catch it: ask the same question two ways and see whether the answer and its sources hold steady.

For companies

The pitch you will hear is that an insight assistant finally puts your whole research archive to work. The honest version is that it lowers the cost of asking your archive a question, which is genuinely useful when the alternative is that nobody asks at all. The risk is that a fluent answer feels like a finished one. A synthesised paragraph with three footnotes looks authoritative whether or not the underlying studies support it, and the people reading the insight in a decision meeting will not click the citations. When you evaluate a vendor, do not accept a claim that the tool checks its sources as if it were a measured accuracy rate. It describes the design, not a result. Ask what share of your product's answers are fully supported by the sources they cite, measured by whom, on what test. Keep a named human owner for any answer that feeds a real decision.

For political parties

A party's research archive holds years of polling, focus-group transcripts and message tests, and an assistant that can interrogate all of it quickly is also an assistant that can produce a fluent, well-formatted answer that flatters a decision already taken. That is the risk to guard against, especially when the archive is thin on the point and the model fills the gap. The discipline is to check that the answer rests on evidence you actually hold, and to be most suspicious when it tells you what you hoped to hear. An answer engineered, wittingly or not, to confirm the leadership's preference is worthless as intelligence. It costs you the early warning the archive was supposed to provide, and the miss shows up on election night, when it is too late to recount.

For government and policy

Public bodies hold some of the largest research archives anywhere, in evaluations, consultations and statistical reports, and the case for an assistant that can answer across them is strong. It is institutional memory that would otherwise walk out the door with retiring staff. The stakes also make the faithfulness problem sharper. An official answer that cites real reports but misrepresents what they found carries the authority of the source while getting the substance wrong, and a citizen or a minister reading it has every reason to trust it. For public use the defensible posture is to require what the market does not yet volunteer. Insist on traceable citations on every answer, record which questions were answered by the tool and how they were checked, and write your own check into the procurement because there is nothing external to lean on.

How to use this

Before you trust an archive assistant with anything that matters, put it through four questions. First, can you click every claim back to a source, and when you do, does the source actually say what the answer says it does? An answer with no citations, or citations that do not hold, is not a finding. Second, is the archive likely to contain the answer at all, or are you watching the model paper over a gap it could not retrieve? Ask the question a second way and see whether the answer survives. Third, does the tool flag stale or conflicting sources, and have you checked that the flag works rather than trusting that it does? Fourth, when a vendor offers a number for accuracy or sources checked, can that number be traced to a real test, or is it a description of intent? The capability is a genuine advance. What has not arrived is a reason to skip reading the sources, and the day a smooth answer persuades you to stop checking is the day the tool starts costing you more than it saves.

Case studies

Stravito Assistant (2024 onward). Stravito's Assistant is the clearest published example of the archive-that-answers pattern. A user asks a question in natural language, and the system returns a synthesised answer with page-level references back into the owned research library, under a framing the company calls Glass Box AI to stress that every claim is traceable (Stravito, 2024). Lavazza Group is a named adopter, having built an internal assistant on the platform that drew industry recognition in 2026 (Stravito, 2026).

Two caveats have to travel with the example. Everything known about how well it works is the vendor's own account, and no independent faithfulness rate has been published. The value of the design rests entirely on the reader using the page-level references it provides, rather than trusting the fluent paragraph above them. The tool makes checking possible. It does not do the checking for you.

References

Brehme, L., Dornauer, D., Ströhle, T., Ehrhart, T. and Breu, R. (2025) 'Retrieval-augmented generation in industry: an interview study on use cases, requirements, challenges, and evaluation', arXiv preprint arXiv:2508.14066. Available at: https://doi.org/10.48550/arXiv.2508.14066 (Accessed: 18 August 2026).

Bruckhaus, T. (2024) 'RAG does not work for enterprises', arXiv preprint arXiv:2406.04369. Available at: https://doi.org/10.48550/arXiv.2406.04369 (Accessed: 18 August 2026).

Choubey, P.K., Peng, X., Bhagavath, S.K.S., Huang, K.-H., Xiong, C. and Wu, C.-S. (2025) 'Benchmarking deep search over heterogeneous enterprise data', arXiv preprint arXiv:2506.23139. Available at: https://doi.org/10.48550/arXiv.2506.23139 (Accessed: 18 August 2026).

Dovetail (2024) Ask Dovetail. Available at: https://dovetail.com/product/ask-dovetail/ (Accessed: 18 August 2026).

Es, S., James, J., Espinosa-Anke, L. and Schockaert, S. (2024) 'RAGAS: automated evaluation of retrieval augmented generation', Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, pp. 150–158. Available at: https://doi.org/10.18653/v1/2024.eacl-demo.16 (Accessed: 18 August 2026).

Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., Wang, M. and Wang, H. (2023) 'Retrieval-augmented generation for large language models: a survey', arXiv preprint arXiv:2312.10997. Available at: https://doi.org/10.48550/arXiv.2312.10997 (Accessed: 18 August 2026).

Glean (2024) Retrieval-augmented generation (RAG): the key to enabling generative AI for the enterprise. Available at: https://www.glean.com/blog/retrieval-augmented-generation-rag-the-key-to-enabling-generative-ai-for-the-enterprise (Accessed: 18 August 2026).

Karakurt, E. and Akbulut, A. (2025) 'Retrieval-augmented generation (RAG) and large language models (LLMs) for enterprise knowledge management and document automation: a systematic literature review', Applied Sciences, 16(1), 368. Available at: https://doi.org/10.3390/app16010368 (Accessed: 18 August 2026).

Liu, N.F., Zhang, T. and Liang, P. (2023) 'Evaluating verifiability in generative search engines', Findings of the Association for Computational Linguistics: EMNLP 2023, pp. 7001–7025. Available at: https://doi.org/10.18653/v1/2023.findings-emnlp.467 (Accessed: 18 August 2026).

Market Logic Software (2023) Market Logic launches DeepSights, the world's first AI assistant for market research and intelligence. Available at: https://marketlogicsoftware.com/news/2023/deepsights-press-release/ (Accessed: 18 August 2026).

Stravito (2024) Stravito Assistant. Available at: https://www.stravito.com/ai-assistant (Accessed: 18 August 2026).

Stravito (2026) How Lavazza Group built AI personas that tell the truth — and why Ad Age took notice. Available at: https://www.stravito.com/resources/how-lavazza-group-built-ai-personas-that-tell-the-truth-and-why-ad-age-took-notice (Accessed: 18 August 2026).

Vectara (2026) Hallucination leaderboard (HHEM). Available at: https://github.com/vectara/hallucination-leaderboard (Accessed: 18 August 2026).

Explore the idea

Let’s talk

Invisible forces shape your world — until you hire Latenta®

Contact