A&A INSIGHTS
Productizing repeat service work: judge by pass rate on hard cases
Whether to productize a repeating request is decided by the pass rate on a set of your own hardest past cases, not by how many engagements you have done.
日本語で読む
THE STARTING POINT
However often the same request arrives, the number of engagements is not grounds for productizing it. Whether to release a fixed offer is decided by first assembling a representative set of cases across customers — the ones that overran the estimate, the ones that needed rework and the ones you declined — and then reading its pass rate.
Judge productization by pass rate on representative cases, not by case count
A&A perspective
Whether a repeating request should become a fixed offer is not decided by how many times you have taken that request. It is decided by first assembling a representative set of hard cases drawn from your own past engagements, and then seeing how many of them pass under one fixed procedure. The reader assumed here is the founder of a contract services business who delivers the work with AI and who is taking similar requests from several different customers, pricing each one separately. A rising case count gives you the feeling that demand exists, but it tells you nothing about whether one fixed procedure and one fixed price survive the variance between those customers. That is the answer; the rest of this article covers how to build the set and how to turn a pass rate into a decision. This is A&A's proposal, not a procedure the sources set out for contract businesses.
From the sources
The Newfront case study published by Anthropic explains that while the company was looking for an AI to power its benefits chatbot, it had compiled a set of twenty critical customer questions that existing AI models struggled to answer. The page states that those twenty concerned specific benefit plans, insurance terminology and complex coverage scenarios. It quotes Gordon Wintrob, Co-founder and CTO, saying that without any changes to their system, Claude correctly handled twelve of their twenty toughest cases. The same page lists the company size as Small, the industry as insurance and the location as North America.
A&A perspective
What a contract business can carry over from this account is not the ratio of twelve to twenty. Two other things carry over: the order and the selection rule. The order is that the yardstick was fixed before the decision to expand. The selection rule is that the yardstick was not a comprehensive list of questions but a collection of questions that had already failed. The ratio itself cannot be carried over. How the twenty were chosen, what counted as correct, who graded them and when, and why the remaining eight failed are not published, and the account is the vendor's and the customer's own. It is not an independent evaluation, so it is not a forecast of anyone else's pass rate.
Fix the price first and you lose the words to decline
From the sources
In "What Is Revenue Architecture?", published on 20 July 2026, GTM Vault's Rick Koleta writes that revenue architecture treats sequencing as the first decision, because the same tools installed in the wrong order produce fragmentation instead of compounding. The essay sets out eight layers and describes the first one, Identity, as defining who the system is for tangibly enough to disqualify in real time. Of the pricing layer it says that pricing is not a number but a structural decision that reshapes qualification, motion and margin at once, and it attaches the law that pricing restructures every layer it touches.
A&A perspective
That essay is the author's argument about the revenue systems of B2B companies that are already past product-market fit; it does not establish that the same ordering applies to a one-person contract business. The layer order nevertheless maps directly onto the decision to productize contract work. Releasing a fixed offer is work on the pricing layer. Before it, you need to be in a position to judge, at the moment an inquiry arrives, whose requests and which requests you will take. That reading is A&A's. The representative-problem set is the material that makes that judgement possible. If you fix the price without that set, the only thing you have fixed is the amount you invoice; what you will not take remains undefined, and out-of-scope requests flow in at the same price.
A&A perspective
Getting the order wrong shows up not as a mispriced number but as having no words with which to decline. While you are still estimating case by case, each estimate absorbs the variance: you put a higher number on a hard engagement, and you can decline an impossible one. A fixed offer is a decision to remove that absorbing layer. What remains after you remove it is a single line between what the offer covers and what it does not, and if that line has not been drawn, the variance moves into your margin and your own hours.
| How the set comes out | Route to take | Decide this first |
|---|---|---|
| Most pass, and the failure reasons converge on one kind of cause | Release it as a fixed offer | Copy the failing cases out as the definition of out of scope |
| The reasons converge but the number of failures is large | Release a fixed offer with a narrower scope | Choose whether to narrow on input format or on frequency |
| The failure reasons are scattered and share no name | Keep estimating case by case | Decide what to record next so the reasons can be named |
| Every case passes but the contract terms differ per customer | Fix the procedure, quote the contract terms separately | Write subcontracting, data location and incident response time separately |
| You still have only a few requests of the same kind | Do not enter the productization judgement | Decide again whose work, and which part of it, you address |
Build the set from four quadrants of your own past engagements
A&A perspective
You select the representative set yourself, from your own past engagements, rather than asking a customer to supply cases. Lay out the engagements in which you took the same kind of request, across customers, and sort them into four quadrants: the ones that overran the estimate, the ones that needed rework after delivery, the ones you declined at the inquiry stage, and the ones that finished on the original estimate. Draw most of the set from the first three quadrants. From the fourth you need only a few cases, to confirm that passing is possible. Around twenty cases is a working size, but that number is matched to the Newfront account rather than a measured optimum. If you only have twelve, start with twelve.
From the sources
What the Newfront case study describes is not a comprehensive question bank but a set of questions that existing models struggled to answer. In the quotation the page carries from Wintrob, that set is referred to as their twenty toughest cases.
A&A perspective
The reason to select from the hard cases is not to confirm that things pass but to learn in advance where they fail. A set made only of engagements that went smoothly will always produce a high pass rate, and that pass rate cannot be used to decide anything. A fixed offer does not break on the average engagement; it breaks on the ones at the edge. Including the engagements you declined matters most. A declined engagement is a record of a line you have already drawn, and it lets you test whether that line can be copied out as a condition of the offer. Note that what is being assembled here is a set of your own engagements across customers. Setting the acceptance conditions of a single engagement from concrete cases is a different purpose, covered in "Choosing your first paid AI service: how a solo founder carves out one sellable unit of work".
Write down what counts as passing before you run the set
A&A perspective
Once the set exists, decide in writing how you will grade it, before you run it. Three things need deciding. First, the state that counts as passing: distinguish between output you can hand to the customer as it is and a draft you intend to review and correct yourself. Second, a time ceiling: a case that exceeds the per-case time you decided on fails even if it was finished, because a fixed offer is a promise that includes a time ceiling. Third, where a failing case goes: write, case by case, whether a person takes it over or whether it is declined as out of scope.
From the sources
The Newfront case study records that twelve cases were handled correctly without any changes to their system. The page also says the company maintains a flexible approach, choosing the right AI for each specific challenge, and quotes Wintrob saying that even with today's AI capabilities they have hundreds of workflows to enhance.
A&A perspective
The statement that the measurement was taken with no changes made shows that the yardstick was fixed ahead of the work. Keep the same order in contract work. If you look at the pass rate and then loosen the grading criteria, that pass rate stops being grounds for productizing and becomes a procedure for convincing yourself. The practical benefit of working this way is that when you change the procedure or the model, you can measure again against the same set. This is A&A's proposed practice, however; the Newfront case study sets out neither a grading rule nor a re-measurement practice.
Turn the pass rate into one of three routes

A&A perspective
Do not treat the pass rate as a score; use it to choose between three routes. If most cases pass and the reasons the failures failed converge on the same kind of cause, you are in a state where a fixed offer is reasonable. If the reasons converge but the number of failures is large, fix the offer with a narrower scope: limit the input to a single format, or put a ceiling on revision frequency. If the reasons are scattered, keep estimating case by case even when the pass rate looks high, because scattered reasons mean the work is not yet one procedure that can be described. The table below summarises the three. No numeric threshold is given, because the appropriate level moves with the nature of the work and the gross margin per case, and because no measured threshold exists.
From the sources
The same essay sets out a corrective sequence of three steps: eliminate the manual data hops first, because every agent downstream reasons over that data; fix qualification before scaling volume, because volume against a broken filter produces expensive noise; and install agents on clean structure, not instead of it. It also writes that the same eight tools produce compounding growth in one company and expensive chaos in another.
Hypothetical example
Consider a hypothetical. A contract development firm has, over eight months, taken requests from five companies for an internal-rules Q&A assistant covering company regulations and manuals, and has estimated each one separately. This is an invented setting for explanation, not a real engagement. Sorting the past engagements into the four quadrants, the one that overran was the engagement whose regulations existed only as paper scans; the one that needed rework was the engagement whose documents were revised several times a month; the one that was declined contained questions close to labour and HR judgements. The failures cluster around input format and revision frequency, while the labour questions are a different kind of cause altogether. What that layout shows is the scope of a possible fixed offer: limit it to electronic documents revised at most once a month, and place questions touching labour and HR judgement outside the offer.
Turn the out-of-scope half into one sentence you can say at inquiry time
From the sources
The Newfront case study notes that many insurance-tech companies try to solve the industry's inefficiencies by fully automating insurance, believing that removing human brokers will reduce costs and speed up processes, and then presents the company's own position as a different one. It quotes Wintrob saying the insurance industry needs both human expertise and advanced technology, and that their platform uses AI to handle routine tasks so their brokers can focus on providing strategic guidance and solving complex problems for clients.
A&A perspective
What is instructive here is that the part the company does not take on is declared from the start rather than discovered later. A fixed offer needs the same declaration. The cases that failed your representative set are the definition of what is out of scope. Turn it into a sentence like this: name, as a noun, the condition the failing cases share; state what you do when a request meets that condition; and keep both short enough to paste into a reply to an inquiry. It should be something you can say in your first reply, not a small note at the bottom of a price list.
Hypothetical example
For the hypothetical firm above, the sentence would read something like this: "This offer covers company regulations maintained in Word or PDF and revised at most once a month. It does not cover cases where the only original is a paper scan, or cases that require answering questions involving labour and HR judgement. For the former we quote the digitisation step separately; for the latter we take the work as a design that routes the question to the responsible person inside your company." The point is that the out-of-scope half does not end in a refusal: one alternative way of taking the work is attached. The formats and the frequency named in that sentence are provisional values matched to the invented setting; your actual conditions come from your own failing cases.
Where this judgement does not apply, and the next step
From the sources
The essay also states its own boundary: companies without product-market fit do not need revenue architecture, they need a product people want. It adds that architecture compounds a working motion and cannot create one, and that installing it early is how teams end up with elegant systems for selling something nobody buys.
A&A perspective
That boundary transfers to contract work. While the requests are not actually repeating — while you have a few engagements that merely look alike — there is nothing for a representative set to measure. What is needed at that stage is not a productization judgement but deciding again whose work, and which part of it, you are addressing. There is also a case where the judgement passes and a fixed offer is still wrong: when the variance between customers lives not in the work itself but in how contractual responsibility is placed. Conditions such as whether subcontracting is permitted, where data may be stored and the response time required during an incident do not fail a representative set. They appear as every case passing while the same contract terms remain unacceptable. In that situation, keep the procedure fixed and quote the contractual conditions separately.
A&A perspective
Finally, the quadrant construction, the twenty-case guideline and the three routes set out here are all designs A&A proposes, not measured methods, and they are not presented as A&A's own productization record or as customer results. If you would rather see the whole workflow first, read "AI-native GTM: a practical guide for solo founders and small teams"; if you want to compare the unit of work different companies sell, read "AI-native service-company cases: what Advolve, Newfront and Lindy help founders compare". If your procedure and your out-of-scope line are already written and what remains is building it, that is a conversation about scoped development. If you cannot yet explain why your failing cases failed, a conversation that starts from deciding what to measure will serve you better.
Whether to productize a repeating request is decided by the pass rate on a representative set of cases, not by the number of engagements. Assemble the set from your own past work across customers, drawing mainly on the engagements that overran, needed rework or were declined. Decide the grading rule in writing before you run it, and use the pass rate to choose between three routes rather than as a score. Copy the failing cases out into a single out-of-scope sentence you can say when an inquiry arrives. Newfront's twelve out of twenty is the vendor's account, not a target pass rate, and the quadrants and the three routes are A&A's proposal rather than a measured method.
Sources & editorial note
Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.
- Newfront modernizes insurance experiences with Claude
Anthropic · Publication/update date not stated on the inspected page
Accessed 2026-10-01 - What Is Revenue Architecture? Definition and 8 Layers
GTM Vault (Rick Koleta) · 2026-07-20
Accessed 2026-10-01
AI-assisted editorial production
A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.
Editorial check: 2026-10-01