A&A INSIGHTS
Acceptance criteria for AI deliverables that vary from run to run
Move the reproducibility promise from identical output to identical judging criteria: who judges, against what standard, over how many runs. Based on two Anthropic engineering posts.
日本語で読む
THE STARTING POINT
What you can promise is identical judging criteria: who looks at what, against which standard, to declare a pass, and how many runs decide it — all written down before work starts. You cannot promise reproducibility in an AI deliverable by fixing the output. The moment a contract says "the same result every time," that clause is a promise you cannot keep.
What you can promise is not identical output but identical judging criteria
A&A perspective
What you can promise is identical judging criteria: who looks at what, against which standard, to declare a pass, and how many runs decide it — all written down before work starts. You re-ran the job during acceptance testing, and the wording or the classification came out different from the sales demo; sitting down with the contract after that experience, your hand stops, because you cannot write "the same result every time." A&A's conclusion is that it is not that you cannot write it — it is that you must not.
A&A perspective
This substitution does not lower the bar. It converts an assumption contracts commonly carry silently — same input, same output — into an explicit passing condition. Left as an assumption, a wobble in the output leaves nothing to decide whether you or the customer is right, and it gets processed as another round of unpaid rework.
A&A perspective
This article separates what the sources technically state, A&A's translation into delivery contracts, and hypothetical worked examples. Both sources are engineering write-ups about evaluation environments and tool design; neither warrants the legal validity or drafting of any contractual clause under Japanese law. Have the wording of any clause checked by your own counsel.
Anthropic: a contract between deterministic systems and non-deterministic agents
From the sources
Anthropic's tool-design post, published on September 11, 2025, states that in computing, deterministic systems produce the same output every time given identical inputs, while non-deterministic systems — like agents — can generate varied responses even with the same starting conditions.
From the sources
The post continues that writing traditional software means establishing a contract between deterministic systems: a function call such as getWeather("NYC") will always fetch the weather in New York City in exactly the same manner every time it is called. It then positions tools as a new kind of software reflecting a contract between deterministic systems and non-deterministic agents.
From the sources
The post illustrates this: when a user asks whether they should bring an umbrella today, an agent might call the weather tool, answer from general knowledge, or ask a clarifying question about location first. It adds that occasionally an agent might hallucinate, or fail to grasp how to use a tool at all.
A&A perspective
What A&A borrows here is the use of the word contract. The post assumes non-determinism and still says a contract is possible — but what that contract fixes is not the output; it is the boundary and the meaning of the response. Translated into a delivery agreement, what should be fixed is not the text of the deliverable but the boundary: what the deliverable must satisfy before it may be accepted.
| Element of acceptance | Written as identical output | Written as identical judging criteria |
|---|---|---|
| Definition of passing | Matches the output shown in the demo | Satisfies every stated condition |
| Who judges | Frequently left unwritten | Name or role, plus the material they review |
| Number of runs | One, implicitly | Representative count x runs, and passes required |
| Customer re-runs and gets different output | Becomes a defect claim, fixed for free | Stated in advance as no defect if the standard is met |
| Requests the standard cannot judge | Negotiated on the spot | Listed in advance as out of scope |
| Record of the judgment | Not retained | Who keeps inputs, outputs and records, and for how long |
| Differences between artifact types | One clause covers everything | Separate granularity for extraction and for generation |
Three eval terms: task, trial and grader become three columns of acceptance
From the sources
Anthropic's evaluation post, published on January 9, 2026, defines its terms explicitly. A task is a single test with defined inputs and success criteria, and each attempt at a task is a trial. The post states: "Because model outputs vary between runs, we run multiple trials to produce more consistent results."
From the sources
The post defines a grader as logic that scores some aspect of the agent's performance, noting that one task can have multiple graders. It further defines the outcome as the final state in the environment at the end of the trial, with the example that a flight-booking agent might say the flight has been booked, but the outcome is whether a reservation exists in the environment's database.
From the sources
The post also states what makes a good task: one where two domain experts would independently reach the same pass/fail verdict. It adds that ambiguity in task specifications becomes noise in the metrics.
A&A perspective
In A&A's translation, these three terms become three columns of your acceptance criteria. The task is a representative example of what the customer actually sends; the trial is how many runs you will make; the grader is who looks at what and declares pass or fail. The definition of outcome carries the most weight. Write not that the deliverable says it is done, but what must exist in the customer's own records for the work to be complete — a definition that transfers directly into acceptance wording rather than staying technical.
Why "every time" is an unkeepable promise: the probability that all k runs pass
From the sources
The same evaluation post presents pass^k as a consistency metric, stating that pass^k measures the probability that all k trials succeed, and that as k increases pass^k falls, since demanding consistency across more trials is a harder bar to clear. Its worked figure: an agent with a 75% per-trial success rate run three times has roughly a 42% probability of passing all three, since 0.75 cubed is about 0.42.
From the sources
The post notes that this metric matters especially for customer-facing agents, "where users expect reliable behavior every time." It also observes that a task that passed on one eval run might fail on the next. It adds that pass@k and pass^k diverge as trials increase: by k=10, "pass@k approaches 100% while pass^k falls to 0%."
A&A perspective
A&A's translation: what pass^k measures is continuing to meet a standard, not continuing to emit the same string. And identity of output is a stronger promise than meeting a standard. If even criteria-passing becomes harder to guarantee as k grows, promising identical wording is a stronger claim still. Writing "the same result every time" into a contract takes on that stronger promise with no ceiling on k. A contract carrying an unkeepable promise looks protective of the customer while actually producing a state in which nobody can determine where a breach begins.
From the sources
The post states the non-independence point itself: if multiple distinct trials fail because of the same limitation in the environment (its example is limited CPU memory), those trials are not independent, because they are affected by the same factor, and the eval results become unreliable for measuring agent performance.
A&A perspective
That carries a caution into delivery work. The 0.75-cubed figure assumes the three trials are independent, and in real delivery an idiosyncrasy in the input or a gap in the source material can make all three fail in the same way. But non-independence does not weaken the conclusion: it changes only how fast — or whether — the probability falls, because a promise of "every time" is broken by a single differing run. These figures also concern measurement inside an evaluation environment and do not necessarily transfer as a method of quality assurance for a deliverable. What is borrowed is the direction — demanding more consistency raises the bar — not the number.
Six lines to put in the acceptance criteria

From the sources
The principle behind this substitution — not failing correct work over wording — is stated outright in the sources, in the evaluation context. The tool-design post advises: "Avoid overly strict verifiers that reject correct responses due to spurious differences like formatting, punctuation, or valid alternative phrasings." The evaluation post lists, among the weaknesses of code-based graders such as string matching, being "brittle to valid variations that don’t match expected patterns exactly"; it adds that "it’s often better to grade what the agent produced, not the path it took," and recommends building in partial credit for tasks with multiple components.
A&A perspective
The six lines A&A proposes are these. (1) Passing standard: what the deliverable must satisfy, written as conditions to be met rather than as the wording of an output. (2) Judge: the person who declares pass or fail, identified by name or role, together with the material they will look at. (3) Number of runs: how many items, run how many times, and how many of them must pass on how many runs to constitute delivery. (4) Treatment of re-runs: whether a different result when the customer re-runs it after acceptance constitutes a defect. (5) Out of scope: the categories of request the passing standard cannot judge. (6) Retained records: who keeps the inputs, outputs and judging records used to decide pass or fail, and for how long.
A&A perspective
The line A&A expects to be most contested is line four. Something passed at acceptance; the customer runs it themselves later and gets a different result. In a contract that does not address this, that single run can become a defect claim. A&A proposes stating explicitly in line four that output differing in wording is not a defect so long as the passing standard is met, and then separately writing the contact point and response window for results that do not meet the passing standard.
A&A perspective
For line two, the evaluation post's test — would two domain experts independently reach the same pass/fail verdict — works as a question to put to yourself. If you and the customer's own contact read your passing standard separately, do you arrive at the same verdict? If not, what you wrote is an impression rather than a standard. A&A's expectation, untested, is that running that check once before signing leaves fewer points to settle in later negotiation.
A hypothetical example: delivering inquiry summaries and classification
Hypothetical example
The following is a hypothetical example, not a real customer and not an A&A result. Suppose you are delivering a process that summarizes a customer's inbound email and assigns it to one of five categories. Ten samples ran cleanly in the sales demo; re-run during acceptance testing, the summary wording changed and one of the ten landed in a different category.
Hypothetical example
Filling in the six lines gives: passing standard — the classification matches the reference category, and the summary contains no fact absent from the original email while including the subject of the request and any requested date. Judge — the customer's intake staff member, looking at pairs of original email and output. Runs — thirty representative items, three times each, with pass/fail set as a count threshold: classification lands in the same correct category on all three runs for at least 28 of the 30, and the summary meets the passing standard on all three runs for at least 28 of the 30. Re-runs — differing summary wording is not a defect so long as the passing standard is met. Out of scope — messages containing several separate requests, and messages that cannot be judged without reading an attachment. Records — inputs, outputs and judging records retained by the delivery side for six months.
A&A perspective
What A&A wants to stress is that this wording does not lower the bar. For classification it sets a condition close to output identity — the same category on all three runs. For the summary it substitutes two conditions (contains nothing false, includes subject and date) for identity of wording. The count threshold follows the source's recommendation to build in partial credit for tasks with multiple components; the level of the threshold itself is a hypothetical figure A&A chose. Even within one deliverable, the promise that can be made differs by the nature of the artifact. Not putting extraction and generation under the same acceptance clause is the first branch you hit when actually using these six lines.
Where this translation breaks down, and what is not being claimed
A&A perspective
Both sources are engineering write-ups published by Anthropic, describing evaluation environments and tool design in technical terms. Neither says anything about the validity of clauses in Japanese delivery contracts, the treatment of defect or non-conformity liability, or applicability to consumer contracts. The six lines here are a workflow design proposal, not a contract template; have the appropriateness of any clause checked by your own counsel.
A&A perspective
The statement that more trials produce more consistent results concerns measurement within an evaluation. It does not necessarily transfer as a quality-assurance method for a deliverable, and the translation into delivery work is A&A's hypothesis. The pass^k worked figure also assumes independent trials, and when the same gap in the source material causes the same failure on every run that figure does not apply — though the direction of the conclusion does.
A&A perspective
What can be promised as reproducibility varies sharply by artifact type. The evaluation post itself varies the form of grading by type: for tasks with objectively correct answers ("What was Company X’s Q3 revenue?") exact match works, while for something as subjective as research quality it says LLM-based rubrics should be frequently calibrated against expert human judgment. Extracting figures, assigning categories, generating prose and generating images each admit a different granularity of passing standard. This article claims no reduction in disputes, no increase in won work and no percentage reduction in rework from adopting these six lines, and presents no A&A track record or customer-side reproducibility rate.
Next step: on the deal that is stuck, write line two only
A&A perspective
You do not need to fill in all six at once. For the engagement currently stuck at the contract stage, write only line two — the judge. Identify who declares pass or fail, and write what that person needs to look at in order to judge. An engagement that proceeds without this risks the judge appearing after delivery, with the standard invented retroactively. Once it is written, add one line of passing standard that the same person would read and arrive at the same verdict on.
A&A perspective
For separating the demo from acceptance during a paid validation stage, see "What to promise in a paid AI pilot: separating the demo from the acceptance conditions". For how much human checking constrains your capacity once the passing standard exists, see "Before taking on more AI delivery work, estimate where human review jams". The overall picture is in "AI-native GTM: a practical guide for solo founders and small teams". If the target workflow is defined and you want the judging and record-keeping mechanism built, contact us about a scoped development engagement.
You cannot promise reproducibility in an AI deliverable by fixing the output. Anthropic's tool-design post states that deterministic systems return the same output for the same input while agents can generate varied responses from the same starting conditions — and still calls a tool a contract. Its evaluation post defines a task as a single test with defined inputs and success criteria, says multiple trials are run because outputs vary between runs, and shows that the probability of passing all k runs falls as you demand more consistency. So what goes into the contract is: who looks at what against which standard to declare a pass, how many runs, how a re-run with different wording is treated, and what is out of scope. Have the legal appropriateness of any clause checked by your own counsel.
Sources & editorial note
Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.
- Writing effective tools for agents
Anthropic · 2025-09-11
Accessed 2026-10-02 - Demystifying evals for AI agents
Anthropic · 2026-01-09
Accessed 2026-10-02
AI-assisted editorial production
A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.
Editorial check: 2026-10-02