Skip to main content
Navigation

A&A INSIGHTS

Business & AI strategyFor business owners

Measuring delivered-AI quality: acceptance rate instead of an eval set

Stop checking a delivered AI only when a complaint arrives. Spellbook calls acceptance rate better than an eval; Supermetrics contrasts. How to count it and contract it.

after deliveryquality measurementacceptance rateeval setssolo and small service firms
日本語で読む
Scattered documents become an organized comparison and a decision
A conceptual illustration of gathering information, organizing it, comparing conditions and making a decision. Illustration generated with AI

THE STARTING POINT

After you deliver a suggestion-type AI, you can measure quality continuously by how many of its suggestions actually get accepted. But because acceptance rate measures preference rather than correctness, it belongs in the contract as the signal that decides where to look, not in front of the customer as proof of quality.

The answer: acceptance rate is not proof of quality, it is the signal that decides where to look

A&A perspective

For a suggestion-type AI, one that proposes replies, drafts or edits, you can measure quality after delivery by how many of the suggestions it issues actually get accepted. It is cheaper than rebuilding an eval set, and it uses the production suggestions that already flow every day. But acceptance rate does not measure correctness. It measures that customer’s preferences. So it does not belong in front of the customer as proof of quality; it belongs in the contract as the signal that decides which slice to investigate.

A&A perspective

Three things have to be settled to write this into a contract. First, what counts as one: which of accepted, rejected and edited sits in the numerator and the denominator. Second, at what grain you read it: a single overall rate will not narrow down where to look, so you need it by task type or by individual user. Third, what the customer sees: the rate itself as a committed number, or only the report of what you found after investigating a slice that dropped.

A&A perspective

The artifact, stated up front: not a quality-report template. A single table with one row per thing you can measure, giving what it tells you and what it does not. The third column is the point. The accident of turning acceptance rate into a committed number starts with not writing that column.

Why eval sets stop running after delivery

A&A perspective

Acceptance passed. The eval set you built for the sales conversation worked at the time. The problem is what comes after. Input patterns shift, the staff rotate, and you swap the model. The hours to rebuild an eval set in step with that do not exist in a one-to-few-person service firm. The result is that the only trigger for looking at quality becomes a complaint.

A&A perspective

Make complaints the trigger and you always find out after the customer does. And "no complaints" cannot be put in front of a customer as proof of quality: an incident-free stretch is an absence, not a record, and an absence is not material you can explain anything with. So what you need is a signal that arrives daily without being rebuilt. Building a comparable eval set from the buyer’s own task before the sale is covered in "Put the buyer's task in your AI demo: build a comparable eval set"; this article is about what comes after, measurement that uses the production traffic flowing every day once it is delivered.

What can be measured after delivery, and what it does not tell you (A&A’s design. Neither source states how to cut these rows, that edits belong in a third category, or that low-count slices should not be converted to rates. Read only the second column and put it in a contract number and the misreading enters the design.)
What you measureWhat it tells youWhat it does not tell you
Acceptance rate (accepted ÷ issued)Whether it matches that customer’s preferencesCorrectness. An accepted wrong suggestion counts as one too
Share edited before useWhich slice has the direction right and the details wrongHow heavy the edit was. A one-character fix and a rewrite count the same
Breakdown of rejection reasons (optional field)A lead on what caused the dropRejections with no reason entered. The more optional, the more skewed
Consistency across repeated runs of one inputWhether the same judgment comes out each timeCorrectness. Being wrong the same way every time also scores as consistent
Acceptance rate per slice (task type, individual user)Where to look by eye nextThe reality of low-count slices. Smaller denominators appear to move more
Length of the stretch with no complaintsOnly the fact that nothing has been reportedQuality. An incident-free stretch is an absence, not a record
Growth in usage and active usersHow far adoption has spreadQuality. Do not read an adoption metric as a quality metric

The source: Spellbook calls acceptance rate better than an eval

From the sources

On the Spellbook case study page published by Anthropic, Scott Stevenson, CEO and co-founder, states how the company measures: "For every suggestion we provide, we measure how many actually get accepted by a user." He continues: "We think that's better than an eval, because it measures the subjective preferences of the user and whether we're meeting them." The reason given for preferring it to an eval is that it captures the user’s subjective preferences and whether the company is meeting them.

Anthropic ↗

From the sources

The same page records the conditions under which that metric works. The company’s Claude-powered review agents run across about 530,000 contracts every month, and on top of that it handles more than 700,000 chat messages from practicing lawyers each month. The company size field reads "Company size: Medium", and since launching in 2022 it has grown to 5,000 customers across 80 countries with 250 employees.

Anthropic ↗

From the sources

The page also states that chat answers come with "citations a lawyer can check". On the review standards side, it explains that each legal team codifies how it reviews agreements and the company’s agents then apply that standard to every contract that follows.

Anthropic ↗

From the sources

Importantly, the same page states that the company also runs an evaluation pipeline. It grades production models against Claude Fable 5, "the oracle model in its evaluation pipeline", with criteria vetted by legal experts. To decide which model goes where, it says, the team measures candidates against Fable and does manual review with its legal engineers and domain experts. Acceptance rate and the evaluation pipeline are described side by side on the same page.

Anthropic ↗

A&A perspective

First, one reading: A&A takes this to mean acceptance rate has not replaced the company’s evals but runs alongside them. With that in view, put these together and the conditions under which acceptance rate functions become visible: the user can decide accept or reject on the spot, the material for that decision is attached to the suggestion, and there are enough decisions. The source does not present "citations a lawyer can check" as the material for an acceptance judgment; reading verifiability as what makes acceptance rate a meaningful signal is A&A’s. And A&A reads the third condition as the one most likely to be missing in a service engagement. This article substitutes acceptance rate for an eval set because of the provider’s own constraint, no hours to maintain one, not because the company in the source does so.

The source: the company uses a repeat-run consistency check to select models

From the sources

As a check used to decide which model goes where, the page records: "One check runs the same contract review, same prompt and contract ten times in a row and scores how consistently issues get flagged, because a lawyer only trusts a review they don't have to repeat." The same contract and the same prompt, ten times in a row, scored on how consistently issues get flagged, with the stated reason that a lawyer only trusts a review they do not have to repeat.

Anthropic ↗

From the sources

The page also records customer wording as part of the history of a model change. Earlier this year the company moved the review surface from Sonnet 4.6 to Opus 4.6, after customers told the team, in effect: "I feel like I have to run it three or four times to find everything". The page does not say the change was made instead of using a metric.

Anthropic ↗

A&A perspective

A&A reads those two as the same thing from both sides. "I feel like I have to run it three or four times" is the felt expression of low consistency, so a check that converts the feeling back into a number by running one input ten times sits alongside it. Acceptance rate measures preference; this measures stability. But the source records this as a model-selection check, not as post-delivery monitoring; repurposing it after delivery, and holding both, is A&A’s design. The repeat-run mechanism itself is covered for pre-sale evaluation in "Put the buyer's task in your AI demo: build a comparable eval set"; what is new here is running it after delivery, against the slice where rejections rose.

Acceptance rate measures preference, not correctness

A&A perspective

The advantage the source claims for acceptance rate is that it measures the user’s subjective preferences. Turned around, that means it does not measure correctness. Users accept wrong suggestions, and they reject right ones on preference. Acceptance rate does not distinguish either case.

A&A perspective

So using acceptance rate as proof of quality in front of a customer means committing to something you cannot deliver. At 90% acceptance, the possibility that errors sit inside the accepted set is untouched. And when the rate falls, the cause may be degradation, or it may be that a staff member changed and the preferences changed with them.

A&A perspective

It still has a virtue an eval set does not: it arrives every day without being rebuilt, and it picks up something an eval set cannot encode in principle, namely that customer’s own preferences. The place to use it is not proving quality but narrowing down which slice to go and look at. A human looks only at the slice where the rate fell. In that order, you locate degradation without paying to rebuild an eval set.

A&A perspective

Treat the denominator with care. The company in the source operates at about 530,000 contracts and more than 700,000 chat messages a month. In a one-to-few-person service firm, where a given slice might see twenty suggestions in a month, the same stability is not available. A&A’s caution: do not convert low-count slices into rates; list accepted, rejected and edited as raw counts. The smaller the denominator, the more a rate appears to move.

The contrast: a company that does not headline acceptance rate secures quality another way

From the sources

The Supermetrics case study page published by Anthropic headlines a different metric: "Grew active users of its Claude connector 250% month over month on average since its February launch." Average 250% month-over-month growth in active users of its Claude connector since the February launch, that is adoption, not acceptance.

Anthropic ↗

From the sources

The page also states the company put accuracy and consistency first in every round of model testing. Beyond that it accounts for quality by other means. On how data is returned it says the company "Serves normalized data from hundreds of marketing sources, with source and timestamp attached to every figure", so every figure carries a source and a timestamp. On write actions, it explains that campaigns get created in a paused state, not live, so nothing goes to market without a human actively choosing.

Anthropic ↗

From the sources

Pre-rollout validation is described as having been done by the customer. Morten Kleven, a digital marketing strategist at Layer, a 25-person Norwegian agency and Supermetrics customer, spent weeks validating the output before rolling it out to clients and said: "I have not found a single hallucination or incorrect data point." The page adds that client reports that took 10 hours of manual spreadsheet work now take 20 minutes, and that 17 of those minutes go into dropping the finished data into the client’s Google Sheet.

Anthropic ↗

A&A perspective

A&A reads the contrast as a difference of role, not of merit. What the Supermetrics page shows is a configuration: an adoption metric out front, quality secured on the construction side by attaching a source and timestamp to every figure and by not going live until a human chooses, and pre-rollout confirmation left to the customer’s own validation. Acceptance testing before delivery and continuous measurement after it are different things, and a customer spending weeks validating is not a substitute for a mechanism that measures the days that follow. The points here are not to rank the two companies’ metrics against each other, and not to read 250% as a quality figure.

How to count accepted, rejected and edited, and what to write in the contract

A&A perspective

Settle first what happens when something is edited and then used. Accepted, rejected, or a third category? Neither source page states how to draw this distinction; it is each company’s design decision. A&A recommends a third category, because a slice with many edits and a slice with many rejections call for different next steps. Many edits means the direction is right and the details are not. Many rejections means the direction itself is wrong.

A&A perspective

Settle next who records it. A design that makes users write a reason increases unrecorded rejections and skews the denominator. So make only the three-way choice mandatory, accepted, rejected or edited, and keep the reason optional. Expect the optional field to go unfilled; even unfilled, the three-way breakdown alone narrows where to look.

A&A perspective

What goes in the contract is a procedure, not a target number. "Record monthly counts of accepted, rejected and edited per slice." "When rejections in a slice rise above the previous month, the provider reviews ten suggestions from that slice by eye and reports." "This record is not a quality guarantee; it is used to identify where to check." Those three lines contain no promise you cannot keep. Put a target on the acceptance rate and you invite optimization in the wrong direction: issuing fewer suggestions to hit the number.

A&A perspective

The table below is A&A’s design. Neither source states how to cut these rows, that edits belong in a third category, or that low-count slices should not be converted to rates. The third column is the substance; read only the second and put it into a contract number and you build the misreading this article is trying to prevent.

Hypothetical: how much can be measured when monthly volume is small

Hypothetical example

What follows is a hypothetical setup. It is not a real customer and not something A&A carried out. Suppose you delivered, to a forty-person home-equipment company, a system that drafts replies to inbound email. It issues about 180 drafts a month, and three staff use it. Say the task types split three ways: quotation requests, delivery-date checks, and fault reports.

Hypothetical example

An overall acceptance rate across 180 a month is readable as a rate. But split by slice, fault reports run around twenty a month. Read that slice as a rate and one case moves it five points. So in the same setup, report the whole as a rate and the slices as raw counts: "Fault reports: accepted 11, edited 6, rejected 3." In that shape, a month-on-month comparison shows which box moved.

Hypothetical example

Say that in the third month, rejections in fault reports rise from 3 to 9. This is where you run the procedure in the contract rather than reporting an acceptance rate to the customer: review ten drafts from that slice by eye. Suppose the review shows most rejections were drafts that misjudged the warranty period (this is a hypothetical setup). The cause was not degradation but a revision to the warranty rules in April.

Hypothetical example

In the same setup, add one consistency check. Pick one representative input from the slice where rejections rose and run it five times with the same instruction, watching whether the same judgment comes out each time. If it is wrong the same way every time, the cause is on the side of how the rules are written. If the judgment varies run to run, the cause is on the side of an ambiguous instruction. Which it is changes what you fix (the count is part of the hypothetical and is not taken from the source’s ten).

A&A perspective

The absence of money and hours saved in this hypothetical is deliberate. What the acceptance record directly changes is only when you notice degradation and how wide an area you have to search. How churn or follow-on orders move does not come out of this table, and A&A has not measured any relationship between this design and retention or revenue.

Limits, and one step you can take today

A&A perspective

Five limits. First, both pages are customer stories Anthropic published as adoption examples of its own product, not independent verification. The figures they carry are self-reported by the vendor and the companies concerned. This article does not use any of those figures as a forecast for anyone else.

A&A perspective

Second, acceptance rate does not measure correctness. An accepted wrong suggestion and a rejected right one are both counted as one. So it is not proof of quality to hand a customer. Verifying the result itself is covered in "How to verify the business result after AI says it is done".

A&A perspective

Third, the scale differs. The company in the source handles about 530,000 contracts and more than 700,000 chat messages a month and its company size field reads "Company size: Medium". A one-to-few-person service firm will not get the same statistical stability. The risk of reading a small denominator as a rate is stated in the body as A&A’s caution.

A&A perspective

Fourth, this metric does not apply to AI that is not suggestion-shaped. Where the output does not divide into accepted or rejected, classification, extraction, transcription, the definition of an acceptance rate cannot be constructed at all. In those cases only the consistency side of the measurement is available.

A&A perspective

Fifth, how the accepted, rejected and edited distinction gets recorded depends on how the product is built. Which side an edited-and-used suggestion counts on is each company’s design decision and is stated on neither source page. In a product with no mechanism to record it, the first task is making the three-way choice capturable.

A&A perspective

One step today: pick one suggestion-type system you have delivered and check whether you can count how many of last month’s suggestions were accepted. If you cannot, the first task is not metric design but adding the record of the three-way choice. If you can, split it by task type and list any slice under twenty as raw counts. The whole picture from acquisition to retention is in "AI-native GTM: a practical guide for solo founders and small teams".

After you deliver a suggestion-type AI, you can measure quality continuously by how many suggestions get accepted. Its virtues are that it arrives every day without the hours to rebuild an eval set, and that it picks up that customer’s own preferences, which an eval set cannot encode. But because acceptance rate does not measure correctness, it is not proof of quality to hand a customer. Its use is narrowing down which slice to investigate. So what goes in the contract is a procedure rather than a target: review by eye and report on the slice where rejections rose. Record the three-way choice of accepted, rejected and edited, list low-count slices as raw counts rather than rates, and where needed run one input several times to look at consistency separately. What the Spellbook page shows is the company’s stated view that acceptance rate is better than an eval, the fact that it runs alongside the company’s evaluation pipeline rather than replacing it, and the fact that a model-selection check runs one input ten times and scores consistency. What the Supermetrics page shows is a configuration that headlines adoption while securing quality on the construction side, through a source and timestamp on every figure and a human choosing before anything goes live. Both are vendor-published self-reports and demonstrate nothing about reproducibility at small scale. The three-way design, the third category, and the decision not to rate small denominators are A&A’s proposals, and their effect has not been measured.

Sources & editorial note

Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.

  1. Spellbook runs 530,000 contract reviews a month with Claude

    Anthropic · no date shown on page

    Accessed 2026-10-06
  2. Supermetrics lets marketers manage ad campaigns from a conversation with Claude

    Anthropic · no date shown on page

    Accessed 2026-10-06

AI-assisted editorial production

A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.

Editorial check: 2026-10-06

← All articles