Skip to main content
Navigation

A&A INSIGHTS

Business & AI strategyFor business owners

Productizing repeat service work: judge by pass rate on hard cases

Whether to productize a repeating request is decided by the pass rate on a set of your own hardest past cases, not by how many engagements you have done.

productizing servicesfixed offerestimatingAI service deliveryservice businesssolo and small teams
日本語で読む
Scattered documents become an organized comparison and a decision
A conceptual illustration of gathering information, organizing it, comparing conditions and making a decision. Illustration generated with AI

THE STARTING POINT

However often the same request arrives, the number of engagements is not grounds for productizing it. Whether to release a fixed offer is decided by first assembling a representative set of cases across customers — the ones that overran the estimate, the ones that needed rework and the ones you declined — and then reading its pass rate.

Judge productization by pass rate on representative cases, not by case count

A&A perspective

Whether a repeating request should become a fixed offer is not decided by how many times you have taken that request. It is decided by first assembling a representative set of hard cases drawn from your own past engagements, and then seeing how many of them pass under one fixed procedure. The reader assumed here is the founder of a contract services business who delivers the work with AI and who is taking similar requests from several different customers, pricing each one separately. A rising case count gives you the feeling that demand exists, but it tells you nothing about whether one fixed procedure and one fixed price survive the variance between those customers. That is the answer; the rest of this article covers how to build the set and how to turn a pass rate into a decision. This is A&A's proposal, not a procedure the sources set out for contract businesses.

From the sources

The Newfront case study published by Anthropic explains that while the company was looking for an AI to power its benefits chatbot, it had compiled a set of twenty critical customer questions that existing AI models struggled to answer. The page states that those twenty concerned specific benefit plans, insurance terminology and complex coverage scenarios. It quotes Gordon Wintrob, Co-founder and CTO, saying that without any changes to their system, Claude correctly handled twelve of their twenty toughest cases. The same page lists the company size as Small, the industry as insurance and the location as North America.

Anthropic ↗

A&A perspective

What a contract business can carry over from this account is not the ratio of twelve to twenty. Two other things carry over: the order and the selection rule. The order is that the yardstick was fixed before the decision to expand. The selection rule is that the yardstick was not a comprehensive list of questions but a collection of questions that had already failed. The ratio itself cannot be carried over. How the twenty were chosen, what counted as correct, who graded them and when, and why the remaining eight failed are not published, and the account is the vendor's and the customer's own. It is not an independent evaluation, so it is not a forecast of anyone else's pass rate.

Fix the price first and you lose the words to decline

From the sources

In "What Is Revenue Architecture?", published on 20 July 2026, GTM Vault's Rick Koleta writes that revenue architecture treats sequencing as the first decision, because the same tools installed in the wrong order produce fragmentation instead of compounding. The essay sets out eight layers and describes the first one, Identity, as defining who the system is for tangibly enough to disqualify in real time. Of the pricing layer it says that pricing is not a number but a structural decision that reshapes qualification, motion and margin at once, and it attaches the law that pricing restructures every layer it touches.

GTM Vault (Rick Koleta) ↗

A&A perspective

That essay is the author's argument about the revenue systems of B2B companies that are already past product-market fit; it does not establish that the same ordering applies to a one-person contract business. The layer order nevertheless maps directly onto the decision to productize contract work. Releasing a fixed offer is work on the pricing layer. Before it, you need to be in a position to judge, at the moment an inquiry arrives, whose requests and which requests you will take. That reading is A&A's. The representative-problem set is the material that makes that judgement possible. If you fix the price without that set, the only thing you have fixed is the amount you invoice; what you will not take remains undefined, and out-of-scope requests flow in at the same price.

A&A perspective

Getting the order wrong shows up not as a mispriced number but as having no words with which to decline. While you are still estimating case by case, each estimate absorbs the variance: you put a higher number on a hard engagement, and you can decline an impossible one. A fixed offer is a decision to remove that absorbing layer. What remains after you remove it is a single line between what the offer covers and what it does not, and if that line has not been drawn, the variance moves into your margin and your own hours.

Reading the pass rate: three routes, and what to decide before looking
How the set comes outRoute to takeDecide this first
Most pass, and the failure reasons converge on one kind of causeRelease it as a fixed offerCopy the failing cases out as the definition of out of scope
The reasons converge but the number of failures is largeRelease a fixed offer with a narrower scopeChoose whether to narrow on input format or on frequency
The failure reasons are scattered and share no nameKeep estimating case by caseDecide what to record next so the reasons can be named
Every case passes but the contract terms differ per customerFix the procedure, quote the contract terms separatelyWrite subcontracting, data location and incident response time separately
You still have only a few requests of the same kindDo not enter the productization judgementDecide again whose work, and which part of it, you address

Build the set from four quadrants of your own past engagements

A&A perspective

You select the representative set yourself, from your own past engagements, rather than asking a customer to supply cases. Lay out the engagements in which you took the same kind of request, across customers, and sort them into four quadrants: the ones that overran the estimate, the ones that needed rework after delivery, the ones you declined at the inquiry stage, and the ones that finished on the original estimate. Draw most of the set from the first three quadrants. From the fourth you need only a few cases, to confirm that passing is possible. Around twenty cases is a working size, but that number is matched to the Newfront account rather than a measured optimum. If you only have twelve, start with twelve.

From the sources

What the Newfront case study describes is not a comprehensive question bank but a set of questions that existing models struggled to answer. In the quotation the page carries from Wintrob, that set is referred to as their twenty toughest cases.

Anthropic ↗

A&A perspective

The reason to select from the hard cases is not to confirm that things pass but to learn in advance where they fail. A set made only of engagements that went smoothly will always produce a high pass rate, and that pass rate cannot be used to decide anything. A fixed offer does not break on the average engagement; it breaks on the ones at the edge. Including the engagements you declined matters most. A declined engagement is a record of a line you have already drawn, and it lets you test whether that line can be copied out as a condition of the offer. Note that what is being assembled here is a set of your own engagements across customers. Setting the acceptance conditions of a single engagement from concrete cases is a different purpose, covered in "Choosing your first paid AI service: how a solo founder carves out one sellable unit of work".

Write down what counts as passing before you run the set

A&A perspective

Once the set exists, decide in writing how you will grade it, before you run it. Three things need deciding. First, the state that counts as passing: distinguish between output you can hand to the customer as it is and a draft you intend to review and correct yourself. Second, a time ceiling: a case that exceeds the per-case time you decided on fails even if it was finished, because a fixed offer is a promise that includes a time ceiling. Third, where a failing case goes: write, case by case, whether a person takes it over or whether it is declined as out of scope.

From the sources

The Newfront case study records that twelve cases were handled correctly without any changes to their system. The page also says the company maintains a flexible approach, choosing the right AI for each specific challenge, and quotes Wintrob saying that even with today's AI capabilities they have hundreds of workflows to enhance.

Anthropic ↗

A&A perspective

The statement that the measurement was taken with no changes made shows that the yardstick was fixed ahead of the work. Keep the same order in contract work. If you look at the pass rate and then loosen the grading criteria, that pass rate stops being grounds for productizing and becomes a procedure for convincing yourself. The practical benefit of working this way is that when you change the procedure or the model, you can measure again against the same set. This is A&A's proposed practice, however; the Newfront case study sets out neither a grading rule nor a re-measurement practice.

Turn the pass rate into one of three routes

A flow diagram running left to right in one direction. Stage one is "repeating requests": similar requests arriving from several customers, each priced separately. Stage two is the "representative case set": about twenty cases gathered across customers from your own past engagements, drawn mainly from the ones that overran the estimate, needed rework, or were declined. Between stage two and stage three a vertical human-judgement gate reads "write the grading rule first", marking the three things decided before the set is run: the state that counts as passing, the per-case time ceiling, and where a failing case goes. Stage three is the "pass rate", treated as material for a decision rather than a score. Stage four is the "three routes": release a fixed offer, release a narrower fixed offer, or keep estimating case by case. Pricing is only settled at stage four, at the end of the sequence, and skipping the gate leaves the pass rate unusable as material for the decision.

A&A perspective

Do not treat the pass rate as a score; use it to choose between three routes. If most cases pass and the reasons the failures failed converge on the same kind of cause, you are in a state where a fixed offer is reasonable. If the reasons converge but the number of failures is large, fix the offer with a narrower scope: limit the input to a single format, or put a ceiling on revision frequency. If the reasons are scattered, keep estimating case by case even when the pass rate looks high, because scattered reasons mean the work is not yet one procedure that can be described. The table below summarises the three. No numeric threshold is given, because the appropriate level moves with the nature of the work and the gross margin per case, and because no measured threshold exists.

From the sources

The same essay sets out a corrective sequence of three steps: eliminate the manual data hops first, because every agent downstream reasons over that data; fix qualification before scaling volume, because volume against a broken filter produces expensive noise; and install agents on clean structure, not instead of it. It also writes that the same eight tools produce compounding growth in one company and expensive chaos in another.

GTM Vault (Rick Koleta) ↗

Hypothetical example

Consider a hypothetical. A contract development firm has, over eight months, taken requests from five companies for an internal-rules Q&A assistant covering company regulations and manuals, and has estimated each one separately. This is an invented setting for explanation, not a real engagement. Sorting the past engagements into the four quadrants, the one that overran was the engagement whose regulations existed only as paper scans; the one that needed rework was the engagement whose documents were revised several times a month; the one that was declined contained questions close to labour and HR judgements. The failures cluster around input format and revision frequency, while the labour questions are a different kind of cause altogether. What that layout shows is the scope of a possible fixed offer: limit it to electronic documents revised at most once a month, and place questions touching labour and HR judgement outside the offer.

Turn the out-of-scope half into one sentence you can say at inquiry time

From the sources

The Newfront case study notes that many insurance-tech companies try to solve the industry's inefficiencies by fully automating insurance, believing that removing human brokers will reduce costs and speed up processes, and then presents the company's own position as a different one. It quotes Wintrob saying the insurance industry needs both human expertise and advanced technology, and that their platform uses AI to handle routine tasks so their brokers can focus on providing strategic guidance and solving complex problems for clients.

Anthropic ↗

A&A perspective

What is instructive here is that the part the company does not take on is declared from the start rather than discovered later. A fixed offer needs the same declaration. The cases that failed your representative set are the definition of what is out of scope. Turn it into a sentence like this: name, as a noun, the condition the failing cases share; state what you do when a request meets that condition; and keep both short enough to paste into a reply to an inquiry. It should be something you can say in your first reply, not a small note at the bottom of a price list.

Hypothetical example

For the hypothetical firm above, the sentence would read something like this: "This offer covers company regulations maintained in Word or PDF and revised at most once a month. It does not cover cases where the only original is a paper scan, or cases that require answering questions involving labour and HR judgement. For the former we quote the digitisation step separately; for the latter we take the work as a design that routes the question to the responsible person inside your company." The point is that the out-of-scope half does not end in a refusal: one alternative way of taking the work is attached. The formats and the frequency named in that sentence are provisional values matched to the invented setting; your actual conditions come from your own failing cases.

Where this judgement does not apply, and the next step

From the sources

The essay also states its own boundary: companies without product-market fit do not need revenue architecture, they need a product people want. It adds that architecture compounds a working motion and cannot create one, and that installing it early is how teams end up with elegant systems for selling something nobody buys.

GTM Vault (Rick Koleta) ↗

A&A perspective

That boundary transfers to contract work. While the requests are not actually repeating — while you have a few engagements that merely look alike — there is nothing for a representative set to measure. What is needed at that stage is not a productization judgement but deciding again whose work, and which part of it, you are addressing. There is also a case where the judgement passes and a fixed offer is still wrong: when the variance between customers lives not in the work itself but in how contractual responsibility is placed. Conditions such as whether subcontracting is permitted, where data may be stored and the response time required during an incident do not fail a representative set. They appear as every case passing while the same contract terms remain unacceptable. In that situation, keep the procedure fixed and quote the contractual conditions separately.

A&A perspective

Finally, the quadrant construction, the twenty-case guideline and the three routes set out here are all designs A&A proposes, not measured methods, and they are not presented as A&A's own productization record or as customer results. If you would rather see the whole workflow first, read "AI-native GTM: a practical guide for solo founders and small teams"; if you want to compare the unit of work different companies sell, read "AI-native service-company cases: what Advolve, Newfront and Lindy help founders compare". If your procedure and your out-of-scope line are already written and what remains is building it, that is a conversation about scoped development. If you cannot yet explain why your failing cases failed, a conversation that starts from deciding what to measure will serve you better.

Whether to productize a repeating request is decided by the pass rate on a representative set of cases, not by the number of engagements. Assemble the set from your own past work across customers, drawing mainly on the engagements that overran, needed rework or were declined. Decide the grading rule in writing before you run it, and use the pass rate to choose between three routes rather than as a score. Copy the failing cases out into a single out-of-scope sentence you can say when an inquiry arrives. Newfront's twelve out of twenty is the vendor's account, not a target pass rate, and the quadrants and the three routes are A&A's proposal rather than a measured method.

Sources & editorial note

Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.

  1. Newfront modernizes insurance experiences with Claude

    Anthropic · Publication/update date not stated on the inspected page

    Accessed 2026-10-01
  2. What Is Revenue Architecture? Definition and 8 Layers

    GTM Vault (Rick Koleta) · 2026-07-20

    Accessed 2026-10-01

AI-assisted editorial production

A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.

Editorial check: 2026-10-01

← All articles