Skip to main content
Navigation

A&A INSIGHTS

Business & AI strategyFor business owners

When numbers disagree in AI-drafted documents: settle the facts first

Turning customer records into reports? EvenUp states a split between settling facts and drafting. Completion conditions for the facts stage, and what it cannot catch.

AI document generationcontradiction detectionpipeline splitsource citationssolo and small service firms
日本語で読む
Scattered documents become an organized comparison and a decision
A conceptual illustration of gathering information, organizing it, comparing conditions and making a decision. Illustration generated with AI

THE STARTING POINT

The contradictions left in a document an AI wrote cannot be fixed in the drafting stage, because drafting does not question the set of facts it was handed. So the place to improve quality is the stage before writing starts, where the facts are settled once, and that stage needs completion conditions of its own.

The answer: drafting cannot fix a contradiction, so split off a stage that settles the facts

A&A perspective

When numbers or dates disagree inside a document an AI wrote, the place to fix it is not the generation side. Drafting does not question the set of facts it was handed. Even when two statements about the same fact conflict, drafting picks one and writes it up smoothly. So the place to improve quality is the stage before writing starts, where the facts are settled once, and that stage needs completion conditions separate from the drafting stage.

A&A perspective

Three completion conditions belong on the facts stage. First, every extracted fact carries the page number of its source. Second, multiple statements of the same fact have been merged into one. Third, unresolved disagreements number zero, or are explicitly listed as held open. Until those three hold, drafting does not begin. Break the order and the job of a human rereading the whole document comes back.

A&A perspective

The artifact, stated up front: not an improved prompt. A single table listing the stages down the side, with each stage’s completion condition and the errors that stage cannot catch across the top. The third column is the point. Without writing down which errors no stage can catch, you cannot decide what to expect from the human check at the end.

Downstream does not catch an upstream disagreement

A&A perspective

A system that turns records into documents usually gets built as one flow: read it, pull the facts, write it. Getting that far and working is fast. The problem is that when numbers or dates disagree inside the finished document, there is no way to trace afterwards where they diverged. Because you cannot trace it, a human rereads the whole thing, and the time you thought you saved comes back.

From the sources

The EvenUp case study page published by Anthropic states that premise in one sentence: "Nothing downstream would catch two parts of the system disagreeing about the same fact, and one contradiction is enough for an adjuster to discount the entire demand." Downstream does not detect two parts of the system disagreeing about the same fact. And in this domain, the page says, a single contradiction is enough reason for the opposing adjuster to discount the entire claim.

Anthropic ↗

From the sources

Where the disagreements come from is recorded too. Emre Yamangil, Principal Machine Learning Engineer at EvenUp, describes the manual work before the product: "A paralegal with a highlighter worked through stacks of records that were weeks late and out of order," and continues, "The same visit could be documented three different ways by three different providers." The same visit, documented three ways by three providers.

Anthropic ↗

A&A perspective

A&A reads this as the ground for splitting the pipeline. Multiple statements of the same fact already exist on the input side. And downstream does not detect them. If both of those hold at once, the stage that removes contradictions can only sit before the writing stage. However carefully you craft the prompt, if the set handed to drafting still disagrees with itself, the result does not change. The page does describe one consistency pass downstream: Opus checks the blueprint's answers for consistency across the whole document, just once. That pass asks whether the argument holds together, not whether two extracted facts about the same thing agree, which is why it does not substitute for settling the facts first.

Each stage’s completion condition, and the errors that stage structurally cannot catch (A&A’s design. How the stages are cut, the wording of the conditions and the assignment of uncatchable errors are stated on neither source page. Read from the third column: each error there must be assigned explicitly to someone downstream, or the customer finds it after delivery.)
StageCompletion condition for that stageErrors that stage cannot catch
Intake and sorting (which kind of document)Every page carries a document typeA wrong type. Surfaces in the next stage as a question the page cannot answer
Reading (climb from the cheapest method up)Every page has been read; unread pages are listedA number read but misread. Invisible until reconciled
Extraction (pull facts out one at a time)Count of facts equals count of facts with a source page numberA fact absent from the material. Can only be treated as missing
Merging duplicates (multiple statements of one fact)The merge rule is written down and was followedA wrong merge judgment, treating two different facts as one
Resolving contradictions (what disagreement remains)Unresolved disagreements are zero, or listed as held openAll statements agreeing while all of them are wrong the same way
Drafting (write from the settled set)Every figure in the body derives from the settled setFacts right but the argument built wrong
Human check and signatureThe responsible person reviewed the argument and signedWhat was overlooked. A signature does not guarantee coverage

The source: EvenUp states it settles once, then writes

From the sources

The page states the ordering explicitly: "Every fact keeps a citation to its source page, duplicates get merged, and the whole record is settled once, before drafting starts." Every fact retains a reference to its source page, duplicates are merged, and the whole record is settled once before drafting starts. It continues: "Only then does drafting begin."

Anthropic ↗

From the sources

The volume of input is given as well. Each case can arrive as 1,000 or more pages drawn from around a dozen providers, plus billing ledgers and police reports. The company size field reads "Company size: Startup", and the page says the firms using it have resolved more than 200,000 cases.

Anthropic ↗

From the sources

The writing stage is itself divided. Per the page, Opus sets the blueprint, deciding what legal rules apply, what the other side will attack and what the throughline of the case is, and then checks those answers for consistency across the whole document, just once. Sonnet writes each section from that blueprint. And the page states: "Every draft still ends with a person: an attorney reads, checks, and signs it."

Anthropic ↗

A&A perspective

What a service firm can take at its own scale is only the ordering and the idea of completion conditions. A&A reads two things as transferable, splitting the settling stage from the writing stage, and defining the settling stage’s completion by source citation, duplicate merging and zero contradictions. The model allocation and the scale do not transfer as they are. The page does not frame any of this as a matter of service estimates or contracts; carrying the split into service delivery is A&A’s proposal.

The source: sort before reading, and climb from the cheapest read upward

From the sources

The page shows there is design inside the facts stage as well. On reading it says the system "reads records on a ladder, climbing only as far as each page requires", reading each page the most affordable way it can and only climbing higher when it has to. The rungs are "real text, then standard OCR, then Claude's vision for the scanned bills and checkbox forms neither one can handle." Real text, then standard OCR, and visual processing only for the scanned bills and checkbox forms that neither can handle.

Anthropic ↗

From the sources

The reason for paying for the top rung is stated: "Vision costs the most to run, and it earns that cost, because a misread number hurts the client directly: the demand asks for too little, or the adjuster catches the error and the firm's credibility goes with it." A misread number harms the client directly, so the cost is earned.

Anthropic ↗

From the sources

There is also a stage before reading: "Every document also gets sorted first, a billing statement, a clinical note, an intake form, so the system only asks each page what it can actually answer." Sort first into billing statement, clinical note or intake form, so each page is only asked what it can actually answer. Yamangil’s advice is also recorded: "Spend your expensive model where you can't verify cheaply".

Anthropic ↗

A&A perspective

A&A reads both of those rungs as transferring directly to service work. Sort first and the question put to each page gets shorter, and the wasted cycles on questions a page cannot answer disappear. Climb the reading ladder from the cheapest rung and the cost falls below reading every page the most expensive way. What does not transfer are the proportions. How much of a customer’s material is scanned differs per customer, and the page does not state that proportion either.

Errors that are easy to catch, and errors hidden by polish

From the sources

The page divides errors into two kinds. Which model gets which task is decided by "where a mistake could hide", with everyday extraction on Sonnet or Haiku and the jobs where polish could hide an error going to Opus. Yamangil is quoted: "If a model gets a date wrong, that's easy to catch," and "Attorneys and staff can check the source page", followed by "If it builds an argument that's subtly wrong, that mistake looks just as confident and polished as a correct one."

Anthropic ↗

From the sources

Carrying settled facts in a form the user can verify appears in another case study too. The Spellbook case study page published by Anthropic states that chat answers come with "citations a lawyer can check" — citations a lawyer can verify personally.

Anthropic ↗

A&A perspective

Put those two together and what the source-page citation is for looks different. It is not only there to present grounds. Being able to return to the source page is what puts a wrong date on the easy-to-catch side. Drop the citation and the same error moves to the hard-to-catch side.

A&A perspective

And the errors on the hard-to-catch side, where the facts are right but the argument is built wrong, cannot be caught by the facts stage’s completion conditions. That is what belongs in the table’s third column. A&A recommends assigning this class of error explicitly to the human check, and giving the reviewer the job of seeing whether the argument holds rather than the job of reconciling numbers. Reconciling numbers can be done on the machine side once the citations exist.

Give the facts stage completion conditions separate from drafting

A&A perspective

Write the completion conditions so that met or not met comes out mechanically, not so that a human felt it looked fine. Concretely, three of them. First, the count of extracted facts equals the count of facts carrying a source page number. One fact without a citation is a fail.

A&A perspective

Second, multiple statements of the same fact have been merged. The unit of merging is decided by the provider and written down: treat as identical when date, amount and party all match, for instance. Leave the merging to judgment without writing the rule and an error that treats two different facts as one hides underneath the appearance of having been merged.

A&A perspective

Third, unresolved disagreements number zero. What matters here is not letting the machine resolve them. Where a disagreement remains, it goes on a list as held open, not resolved, because which one to take is a business judgment. The facts stage ends either with that list empty, or with the held-open list going to the customer.

A&A perspective

The table below is A&A’s design. How the stages are cut, the wording of the completion conditions, and the assignment of errors each stage cannot catch are stated on neither source page. Read it from the third column. The errors written there are structurally uncatchable at that stage, so each one has to be assigned explicitly to someone downstream. Any error left unassigned is one the customer finds after delivery.

Hypothetical: splitting the stages on a subsidy-application deliverable

Hypothetical example

What follows is a hypothetical setup. It is not a real customer and not something A&A carried out. Suppose you are delivering a system that drafts subsidy applications from the financial statements, equipment quotations and existing business plan received from a small company. Say each case runs about 40 pages, mixing PDFs and scanned images.

Hypothetical example

Built as one flow, the finished application wavers between the revenue figure in the financial statements and the one in the business plan. In the same setup, split the stages. First, sorting: tag every page as financial statement, quotation, business plan or other. Then reading: pages with extractable text are read as text, and only the pages without it are processed as images. That already produces a list of pages that could not be read.

Hypothetical example

Next, extraction. Pull out items such as revenue, equipment acquisition cost and headcount one at a time, attaching a source file name and page number to each. Run the completion condition here for the first time. Suppose three items come out with no citation (this is a hypothetical setup). Those three are either absent from the material or sitting on a page that could not be read, so they go on a list as missing and are not handed to the drafting stage.

Hypothetical example

At the merge-and-contradiction stage, suppose revenue differs between the financial statements and the business plan. Against the rule, the accounting periods differ, so these are separate facts and are not treated as identical. Which one goes in which field of the application is a business judgment, so the machine does not choose; it goes on the held-open list. Once the customer’s check clears the hold, the set of facts is settled, and only then does the drafting stage begin.

A&A perspective

The absence of hours saved and money in this hypothetical is deliberate. What this split directly changes is only whether you can locate where a contradiction arose, and what you ask the human check to do. How much the per-case review time shrinks depends on the state of the material, and A&A has not measured any relationship between this design and hours or rates.

Where this does not apply

A&A perspective

First, where settleable facts are not the centre of the deliverable, this split does not carry over as is. For a proposal or an essay whose centre is interpretation and judgement, the definition of a stage that settles the facts once cannot even be constructed. There, the only transferable part is retaining the source citations.

From the sources

The page says a single contradiction is enough reason for the opposing adjuster to discount the entire claim. The domain is personal injury law; the firms it names are in California, Ohio and Oregon, its location field reads "North America", and it describes sections being written to state-by-state rules.

Anthropic ↗

A&A perspective

Second, the penalty for a contradiction differs. What the source covers is a domain where the other side reads looking for contradictions. For a deliverable with no adversarial reader, the penalty for a contradiction is nothing like that large. In an internal report, a contradiction simply gets fixed when found. And the investment in splitting stages can only be recovered within the range the penalty justifies. So decide on the split not by technical difficulty but by what happens when one contradiction surfaces.

A&A perspective

Third, it assumes the input is documents. The cost ladder only works because there is a unit called a page. Where the input is conversation records or tabular data, the same ladder does not form. Fourth, scale. Taken together, the company-size field quoted earlier and the case volume stated alongside it mean the investment in splitting stages is not always justified for a one-person service firm.

Limits, and one step you can take today

A&A perspective

Four limits. First, both pages are customer stories Anthropic published as adoption examples of its own product, not independent verification, and the figures they carry are self-reported by the vendor and the companies concerned. This article adopts none of the outcome figures on those pages as a claim of its own, such as the reduction in working hours or the increase in settlement offers (the cited page title is kept verbatim). Their populations and definitions are not given, and the page itself states the post-reduction time in more than one way.

A&A perspective

Second, meeting the completion conditions does not make the document correct. If every statement across all pages agrees and all of them are wrong the same way, the zero-contradiction condition still passes. The facts stage guarantees the internal consistency of the set only; whether the set matches reality is a separate question.

A&A perspective

Third, splitting stages adds the work of deciding the format handed between them. On a small engagement that overhead can outweigh fixing the contradictions by hand. The criterion is the one in the previous section: what happens when one contradiction surfaces.

A&A perspective

Fourth, A&A has not measured this design’s effect. No prediction is offered that splitting the stages reduces review time, and no record is claimed. What is offered is the fact that the source adopts the split, the fact that the stated reason is that downstream does not catch an upstream disagreement, and a design for completion conditions drawn from those.

A&A perspective

One step today: take the most recent document-generation system you delivered and count what share of the figures appearing in its output still carry a citation back to a source page. If the citations are not there, the first task is not improving the prompt but attaching sources to the extracted facts. Designing the citations into the deliverable itself is covered in "Source Design for AI Research Reports: Deliverables That Trace Back". Verifying the result after completion is in "How to verify the business result after AI says it is done", and the whole picture from acquisition to retention is in "AI-native GTM: a practical guide for solo founders and small teams".

The contradictions left in a document an AI wrote cannot be fixed in the drafting stage, because drafting does not question the set of facts it was handed. So the place to improve quality is the stage before writing starts, where the facts are settled once, and that stage needs completion conditions of its own. Three of them: every extracted fact carries a source page number, multiple statements of the same fact have been merged under a written rule, and unresolved disagreements number zero or are explicitly held open. Disagreements are not resolved by the machine; they go back to business judgment as holds. Reading puts sorting first, then climbs from real text upward, using the expensive rung only on the pages that need it. What the EvenUp page shows is that the company adopts a stage that settles facts once before drafting starts, that the stated reason is that downstream does not catch two parts of the system disagreeing about the same fact, and that errors divide into an easy-to-catch side and a side hidden by polish. What the Spellbook page shows is a design that carries settled facts in a form the user can verify individually. Both are vendor-published self-reports and demonstrate neither reproducibility at small scale nor the transferability of their outcome figures. How the stages are cut, the wording of the completion conditions and the assignment of uncatchable errors are A&A’s design, and the effect has not been measured.

Sources & editorial note

Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.

  1. EvenUp cuts document drafting from 15 hours to 15 minutes with Claude

    Anthropic · no date shown on page

    Accessed 2026-10-06
  2. Spellbook runs 530,000 contract reviews a month with Claude

    Anthropic · no date shown on page

    Accessed 2026-10-06

AI-assisted editorial production

A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.

Editorial check: 2026-10-06

← All articles