A&A INSIGHTS
Source Design for AI Research Reports: Deliverables That Trace Back
The value of an AI research report is verifiability. Separate fact, interpretation, and unconfirmed items so every conclusion traces back to its source.
日本語で読む
THE STARTING POINT
The delivery value of AI research lies not in a polished summary but in the buyer's ability to verify the basis of each judgment. Separate facts, interpretations, and unconfirmed items, attach the primary source URL, access date, and original excerpt to every conclusion, and deliver anything you could not confirm still marked as unconfirmed.
The answer: deliverables whose basis the buyer can verify
A&A perspective
The decision a founder faces over an AI research deliverable narrows to one question: will you separate facts, interpretations, and unconfirmed items, and deliver the work so that every conclusion can be traced back to the original text of a primary source? The answer is to separate them and make it traceable. A readable summary is only part of the requirement. The ability for buyers to verify the basis of each judgment themselves is the very reason they pay for outside research, and this article fixes that structure and its acceptance criteria.
A&A perspective
The motive behind this decision is commercial loss. Even if you deliver a polished summary, the moment a client asks for the basis and you cannot reach the original text, post-delivery verification starts consuming hours. Time spent verifying is a direct cost, and the slower the answer, the more trust erodes. Source design is quality control and, at the same time, part of pricing, because post-delivery labor belongs in the estimate. This framing is an Automate and Augment proposal, not a measurement of how the research industry actually performs.
Why a summary alone is not enough: extracted fragments lose context
From the sources
In its Contextual Retrieval engineering post, Anthropic explains that traditional RAG removes context when encoding information, which often results in the system failing to retrieve the relevant information from the knowledge base. Documents are split into small chunks for retrieval, and a chunk viewed on its own may no longer say what it is actually about.
From the sources
The same post gives the example of a knowledge base of US SEC filings: a chunk reading 'The company's revenue grew by 3% over the previous quarter.' does not, on its own, specify which company or which time period it refers to, making it difficult to retrieve the right information or to use the information effectively.
A&A perspective
The same problem shows up in research reports. A conclusion sentence detached from its source is a fragment that has lost its context: the reader cannot check whose observation it was or when it was made. Fluency of the summary and verifiability of the conclusion are separate properties. Overlaying the retrieval fragmentation problem onto deliverable source design is an Automate and Augment analogy; Anthropic's post says nothing about designing research deliverables.
Research with many conditions runs asynchronously
From the sources
According to Anthropic's customer story, Genspark launched out of stealth in mid-2024 as an AI search product, then steadily expanded into parallel search and asynchronous deep research that would email users a finished analysis after running for several minutes. Co-founder and CTO Kay Zhu said that if the product was positioned as search, ten seconds was too short for the system to do meaningful work.
From the sources
The story includes a vendor-reported usage example: one paying user spent a week using a traditional search engine to compare credit cards against a long list of personal criteria, then handed the same question to Genspark's asynchronous agent and got the same conclusion in 15 minutes. It is presented as a use case where comparison research with many conditions suits asynchronous execution that takes time.
A&A perspective
If research with many conditions takes time, what the buyer pays for is not a fast summary but a conclusion checked against a long list of criteria, together with the means to check that conclusion again later. The week and the 15 minutes are the vendor's statements from one company's customer story, not predictions of your delivery times or efficiency. Applying this mode of execution to a small research service is treated here as an Automate and Augment hypothesis.
Separate every line into fact, interpretation, or unconfirmed
A&A perspective
Automate and Augment proposes three labels. A fact is a statement you actually read and confirmed in a primary source, and it carries a URL and an access date. An interpretation is your reading of confirmed facts, signed so the reader knows whose judgment it is. An unconfirmed item is one whose primary source you could not reach; you leave it marked as it is rather than filling the gap with inference. Never mixing the three is the minimum requirement for the deliverable.
Hypothetical example
Check this with a fictional example. Suppose a client orders pricing research on three competitors. Company A's base price is recorded as a fact you read on its official pricing page. Company B's discount behavior stays unconfirmed if the underlying remark cannot be verified in a primary source. The skew in price bands visible when the three are lined up is labeled an interpretation and explicitly marked as your judgment. The example shows the working format; it is not a real engagement or a result.
A&A perspective
Do not delete unconfirmed items. Removing them makes the report look finished, but the buyer then decides without knowing a gap exists. Leaving items marked unconfirmed makes the un-investigated scope visible as outside the contract, and it becomes material for negotiating the next round of research and additional fees. Prioritize completeness as decision material over the appearance of polish.
A record format that traces from conclusion back to source text
From the sources
Anthropic's method prepends chunk-specific explanatory context to each chunk before building the embeddings and the search index, and the post reports that adding this context reduced the top-20-chunk retrieval failure rate by 49 percent, and by 67 percent when combined with reranking. Attaching context to a fragment makes it easier to find again later.
A&A perspective
Borrowing this idea for deliverables, Automate and Augment proposes a format with five fields per conclusion: the conclusion sentence, one of the three labels, the primary source URL with its access date, a short excerpt of the relevant original passage, and what could not be confirmed. Porting a retrieval-accuracy improvement into deliverable design is our adaptation, and the reported reduction rates guarantee neither the quality of your report nor the correctness of the research.
Hypothetical example
A fictional filled-in row. The conclusion reads 'Company A's base plan is confirmed as monthly billing.' The label is fact. The URL and access date are recorded. The excerpt is the relevant single line from the pricing page. The unconfirmed field says 'annual payment discount rate is not stated on the page.' From this one row alone, the buyer can return to the basis of the conclusion unaided. The content is a placeholder showing the format, not a real price or engagement.
State the limits and set acceptance criteria in advance
A&A perspective
Better retrieval accuracy does not guarantee the truthfulness of individual statements. Being able to reach the original text and the original text being correct are separate problems. Sections where the primary source cannot be read stay marked unconfirmed and are handed over as items the client verifies before making final decisions. A sourced report does not eliminate the buyer's checking work; it makes that work short and reliable.
A&A perspective
The Genspark story is a customer story published by Anthropic, not an independent audit. The ARR figure and the usage examples are the vendor's own statements. Applying a US startup's setup to a small research service is an Automate and Augment hypothesis. Do not infer sales, close rates, search rankings, or anyone's personal track record from this article.
A&A perspective
Acceptance criteria are agreed with the client before delivery. Automate and Augment proposes three: every conclusion row can be traced to its primary source within a stated number of steps; unconfirmed items are presented as a list; and the time verification work takes is included in the estimate. If the contract makes acceptance end with checking the basis, follow-up questions after delivery become a paid scope change. A traceable deliverable is a structure that protects trust and gross margin at the same time.
Source design is not an appendix to the deliverable; it is the product itself. Separate fact, interpretation, and unconfirmed items, and use a format that traces each conclusion back to the original text: the client's verification becomes short, and follow-up questions become a paid scope change. Improvements in retrieval and another company's story inform your judgment, but they guarantee neither the truth of any statement nor your own sales. Build traceability into the contract as an acceptance criterion.
Sources & editorial note
Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.
- Introducing Contextual Retrieval
Anthropic · 2024-09-19
Accessed 2026-09-27 - Genspark's Super Agent orchestrates 150+ tools with Claude
Anthropic · n.d.
Accessed 2026-09-27
AI-assisted editorial production
A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.
Editorial check: 2026-09-27