Skip to main content
Navigation

A&A INSIGHTS

Business & AI strategyFor business owners

What to promise in a paid AI pilot: separating the demo from the acceptance conditions

The demo works, but what the paid test actually promises is undecided. In a paid AI service pilot, the buyer picks the cases, the cases AI should hand to a person also count as passes, and the checker and their hours are written up front. Primary sources from Lovable and Anthropic support ten rows for the proposal.

paid pilotacceptance conditionsAI service businessfirst customersproposalsscoping
日本語で読む
Scattered documents become an organized comparison and a decision
A conceptual illustration of gathering information, organizing it, comparing conditions and making a decision. Illustration generated with AI

THE STARTING POINT

At the point where a demo leads into a proposal for a paid test, a pilot that has not decided which state counts as delivered on the customer's own data cannot produce a verdict at the end. This article argues that a paid pilot promises not a rerun of the successful demo but a deliverable the customer can adjudicate on inputs the customer chose. Anthropic's evaluation article supplies four points read across into proposal rows: draw tasks from real failures, test where a behaviour should and should not occur, build a reference solution, and ask whether two experts reach the same verdict. Source statements, A&A interpretation and hypothetical examples stay distinct.

What a paid pilot promises is not a rerun of the demo that worked

A&A perspective

A proposal written at the moment the buyer says "let's try it for a fee" often contains only a feature list, a duration and a price. Finish a pilot in that state and the last question left is "so, was that a success?" The answer becomes an exchange of impressions, unpaid extra work accumulates, and the decision about a full contract slides again. What a paid pilot promises is not a second performance of the demo that worked. It is handing over a deliverable that the buyer can judge pass or fail using means the buyer already has, on inputs the buyer chose. The line between a demo and a paid pilot is not the number of features or the length of the period. It is who picks the inputs.

From the sources

Anthropic's published case page for Lovable explains that every new Claude release goes through the same evaluation Lovable has run from the start, measuring how often the system hits a wall and produces an app that is broken or is not what the user asked for. The page quotes Anton Osika, co-founder and CEO, and Alexandre Pesant, who leads product. No publication date is shown on the page.

Anthropic

From the sources

The same page says that gate matters because Lovable's users often cannot read the code themselves, so they are trusting the output to work.

Anthropic

A&A perspective

What A&A carries from that passage into a service business is not the frequency of evaluation but the premise behind it. The buyer often cannot read the substance of the deliverable directly. The same holds in service work: a buyer cannot necessarily judge on the spot whether a document or a dataset the AI produced is sound. If that is the situation, pass or fail in a paid pilot has to be written in terms of whether the buyer can judge it with means they already hold, not in terms of how good the output is. A promise nobody can adjudicate is not a test, even after the test period ends.

Anthropic

A&A perspective

There is an opposing position. A first paid engagement exists to establish a commercial relationship, and the more conditions you write, the heavier the psychological load on the buyer and the more deals you lose; take something small first and build trust. That position genuinely holds in some situations, and the closing section deals with it. The argument here is not to add conditions. It is to move the conditions you write from features to adjudication. The number of lines does not grow.

Which side picks the cases used in the test

A&A perspective

In a demo the seller picks the inputs. Choosing examples that work well is natural and not in itself wrong. The problem is continuing to use inputs of the same character once money is involved. Ten cases the seller selected can all pass while the buyer still holds the original question: what happens on my own cases?

From the sources

Anthropic's engineering article "Demystifying evals for AI agents" (published January 9, 2026, by Mikaela Grace, Jeremy Hadfield, Rodrigo Olivares and Jiri De Jonghe) says, at the step of collecting the initial evaluation dataset, that teams delay building evals because they think they need hundreds of tasks, and that in reality 20-50 simple tasks drawn from real failures is a great start.

Anthropic

From the sources

At the next step the article advises beginning with what you already test manually: the behaviours you verify before each release during development and the common tasks end users try. If you are already in production, it says to look at your bug tracker and support queue, and that converting user-reported failures into test cases ensures your suite reflects actual usage.

Anthropic

A&A perspective

What A&A takes from this is not the count but the provenance of the cases. "Real failures", "bug tracker" and "support queue" are all records left on the side that used the system, not assumptions held by the side that built it. Translated into a paid pilot, selection of the cases moves to the buyer. Ask concretely. Not "please give me cases that are going well", but "please pull ten cases from the past six months where the work came back or had to be redone". Handing selection to the other side raises the probability that your own expectations are wrong. That increase is exactly what becomes the buyer's evidence for a decision.

Anthropic

A&A perspective

The figure of 20 to 50 does not transfer as it stands. The source is written about evaluating a product's agent, and a single customer's pilot may yield only six comparable cases. Making the count a target forces you to mix in cases of a different character to reach the number, which dilutes what pass and fail mean. What belongs in the proposal is not a large count but one line: these ten cases were selected by the buyer.

Anthropic

Hypothetical example

A hypothetical example. Suppose you propose, to a company that services industrial machinery, a system that drafts the customer-facing work report from the field technician's phone submission (a photo and a short note). The demo used three cases the founder chose personally, all with clear photographs. In the paid pilot you ask that company's site manager: please give me ten reports from the past six months that had to be rewritten before going to the customer. The ten that come back include ones where the photo is too dark to read the part number, ones where the technician's description contradicts the previous report, and ones describing work outside the quoted scope. Those ten become the subject of the test. This setting is fictional and used only for explanation.

How the decisions change between a demo and a paid pilot (A&A's framing; includes hypothetical examples)
DecisionDemo you showPaid test
Who picks the casesThe seller picks examples that work wellThe buyer picks from their own records
Where cases come fromThings that passed in your own handsReal cases that came back or were redone
Cases AI should not doNot shownListed up front as a kind of pass
Who declares pass or failThe mood in the meetingOne named person on the buyer's side, using an existing procedure
What adjudication usesA live screen walkthroughThe delivered artefact itself
Hours the checking takesAbsorbed into a meeting of under an hourWritten into line items as the buyer's effort
When the model changes mid-wayNobody noticesRe-checked on the same case set, with the burden agreed in advance
The decision afterwards"Looks good"Conditions for continuing and for stopping are written down

Put the cases where AI should not act on the pass side

From the sources

At the step of building balanced problem sets the article says to test both the cases where a behaviour should occur and where it should not. One-sided evals create one-sided optimization, it explains, offering the example that if you only test whether the agent searches when it should, you might end up with an agent that searches for almost everything.

Anthropic

A&A perspective

Moved into a paid service pilot, this means listing in advance, as a kind of pass, the cases where the AI should not process the item but hand it to a person with a reason. This is the row most often missing from a proposal. Run the pilot without it and every case handed to a person is automatically counted as a failure. The buyer reads it as "if a human ends up doing it anyway, there is no point", and cases that in fact behaved exactly as designed turn into the reason the deal is lost.

Anthropic

A&A perspective

Write it beforehand and the same result becomes a pass. The wording takes this shape: for this class of case, the correct behaviour is that the AI does not write the body text and instead returns a list of items to confirm to the responsible person. Then decide what has to accompany that return for it to count as a pass. Which fields could not be confirmed, and which inputs were missing. If both are readable from what comes back, it passes. If it arrives silently with blanks, it fails. That way the act of handing work to a person also has a pass and a fail.

Hypothetical example

In the fictional example above, two cases where the photo is too dark to read the part number and one where the description contradicts the previous report are cases whose correct answer is to return them without drafting. The pass condition is that the unreadable field is named concretely, as in "part number (illegible in photograph)", and that the point to re-ask the technician is written in one sentence. For the remaining seven, a pass means the site manager can send them to the customer without editing. The split of seven and three is an illustrative figure; the real numbers follow from the content of the cases that were selected.

Finish exactly one case yourself before you price it

From the sources

The article says it is useful to create, for each task, a reference solution: a known working output that passes all graders. Having one proves that the task is solvable and verifies that the graders are correctly configured.

Anthropic

A&A perspective

Moved into the proposal sequence, this means that before pricing you finish one of the ten buyer-selected cases yourself and show it. There are two purposes. One is confirming the work is genuinely solvable. The other is confirming that the acceptance wording is actually usable. Finishing one case makes it obvious when a pass condition you already wrote cannot adjudicate anything. "A readable report" cannot adjudicate; until it descends to "the part number and the labour hours are recorded, and the number of photographs corresponds to the work items in the body", the other side cannot state a verdict.

Anthropic

From the sources

The article adds that a good task is one where two domain experts would independently reach the same pass/fail verdict. It asks whether those experts could pass the task themselves, says the task needs refinement if not, and explains that ambiguity in task specifications becomes noise in metrics.

Anthropic

A&A perspective

That sentence works directly as a self-check on the proposal. If two people at the buyer's company, say the site manager and the quality assurance lead, look at the same deliverable and state different verdicts, the acceptance condition is not yet written. The source is talking about noise in metrics; in service work the noise lands on the invoice. A case where the verdict splits usually becomes unpaid rework. Before proposing, have those two people read the condition and ask only "can you adjudicate with this?" That alone shows you which sentence to rewrite.

Anthropic

A&A perspective

This one case is unpaid time, and whether it is worth it is conditional. It is worth it when the target work is routine, comparable cases arise several times a month, and there is room for continuation beyond the pilot. For work that occurs a few times a year, or when the pilot fee is below the cost of doing one case, producing a finished example guarantees a loss. In that situation there is an alternative: instead of building the example, obtain one deliverable the buyer produced in the past and attach it to the condition text as the model of a pass. Your own effort does not increase and the standard for adjudication becomes concrete.

Put the person who checks, and the hours it takes them, in the line items

A&A perspective

Once the pass conditions are written, the next question is who adjudicates and when. Three things go into the proposal. One person who checks, named by role. The checking method, which should be a procedure that person already uses. And the expected hours the checking will take. Pilots that omit the third one stall often.

From the sources

The article defines the outcome of an evaluation as the final state in the environment at the end of the trial. Even if a flight-booking agent says "Your flight has been booked" at the end of the transcript, it explains, the outcome is whether a reservation exists in the environment's SQL database.

Anthropic

A&A perspective

What A&A carries from this into the paid-pilot agreement is writing the check in terms of the downstream state of the deliverable. Not "a draft was generated" but "the site manager sent it to the customer without editing" is the pass. How to design the completion check itself is a separate question, and the implementation side is covered in "How to verify the business result after AI says it is done". What is decided here is whose hours perform that check and how those hours are treated inside the agreement.

Anthropic

A&A perspective

If the buyer has no hours for checking, the pilot stalls for reasons unrelated to the AI's performance. At the proposal stage, ask: who can spend how many hours checking these ten cases? If the answer is that no hours are available, reduce the count. Cutting ten to five and extending the period leaves more result than keeping ten and finishing with nothing adjudicated. A quote issued without asking this has silently put the other side's labour into your budget.

Hypothetical example

Back to the fictional example: the person who checks is the site manager. The method is the existing procedure, where the manager reads the report and signs it before it goes to the customer. The time is fifteen minutes per case, two and a half hours for ten. Those two and a half hours go explicitly into the schedule section of the proposal, down to which days inside the two weeks they are taken. Hand over ten cases during a week the manager is out on site and the checking slips to the following week, and the period ends with no conclusion. All figures here are illustrative.

The paid-pilot agreement on one page

A&A perspective

Put all of the above on one page. Ten rows. The target work. Who selects the cases, and how many. The examples that should pass. The examples that should be handed to a person, and their pass condition. The reference finished example. The person who checks. The checking method. The hours the checking takes. How re-checking is handled when the model or the procedure changes. The condition for continuing after the pilot, and the condition for stopping. While those ten rows cannot be filled, you can decide not to charge yet. The usual reason they cannot be filled is that the target work has not yet been narrowed to one thing.

Hypothetical example

Filled in with the fictional example it reads as follows. The target work is drafting the customer-facing report from the field submission. The cases are ten selected by the site manager from rewrites in the past six months. Seven should pass, meaning they can go to the customer without editing. Three should be handed to a person, with the unreadable field named and the point to re-ask written. The reference example is one of the ten selected cases, produced by us before the proposal, with the manager answering whether it is acceptable. Checking is done by the site manager using the existing signature procedure, fifteen minutes per case, two and a half hours in total, over two weeks. If the model or the procedure changes, the same ten cases are re-checked, and re-checking during the pilot period is at our cost. The condition for continuing is that at least six of the seven pass without edits and the three return correctly. The condition for stopping is that the reasons the failures failed scatter across three or more kinds. Every count and duration here is illustrative.

A&A perspective

The last two rows, the condition for continuing and the condition for stopping, are not only for the buyer. With a stopping condition written down, you can withdraw too. With a page that says "continue if six pass", you can respond to only four passing by taking the causes home and redesigning, instead of sliding into unpaid additional development. A pilot with no stopping condition is an open-ended warranty from the seller's side.

When the model or procedure changes mid-way, who re-checks

From the sources

On Lovable's page, Pesant, who leads product, says they have followed each model update, and explains the reason as each model being better than the previous one. The page also carries his statement that Claude Sonnet 3.5 was the first model that made agents work, and that Claude Opus 4.5 was the next big step change in reliability on long-horizon tasks, unlocking a new class of projects. Anthropic's evaluation article, for its part, says evals get harder to build the longer you wait: early on product requirements naturally translate into test cases, but wait too long and you are reverse-engineering success criteria from a live system.

AnthropicAnthropic

A&A perspective

The same thing happens in service work. Mid-pilot, or after the full contract starts, you raise the model, rewrite the prompt, add preprocessing. Each time, cases that used to pass may stop passing. This is where the pilot's asset pays off. The ten cases the buyer selected, with pass and fail written against them, become the regression set as they stand. Discard them when the pilot ends and you rebuild the standard for adjudication at every subsequent change. As the source says, building it later is harder.

Anthropic

A&A perspective

What has to be decided is where the burden sits. There are three available shapes. One, at your cost, re-checking the same ten cases whenever something changes; this belongs inside a monthly maintenance fee. Two, at the buyer's cost, receiving checking hours at every change; this creates effort on the buyer's side, so it needs agreement in advance. Three, freeze the model and the procedure for the period and push changes into the next contract. Which is right depends on how fast the target work changes, but when nothing is written the reality is that you do all of it for free. It costs one line, so put it in the proposal.

Where this design does not fit, and the next step

A&A perspective

Start with where it does not fit. There is a stage where the buyer's budget is not yet fixed and all the internal approval needs is evidence that the thing is technically possible. Bring these ten rows into that stage and the conversation stops, because you are asking the other side to decide something they cannot decide yet. What is fast at that stage is a short unpaid demo plus one page stating the scope. A paid test works better once a budget exists and an internal owner has been named. Get the order wrong and a careful proposal becomes the reason the deal is lost.

A&A perspective

The limits of the numbers are worth stating too. The source's guide of 20 to 50 is about evaluating a product's agent and is not a count that transfers to a single customer's pilot. If only six comparable cases exist, start with six. The step of building one finished example also does not hold when the pilot fee is below the cost of the work. The legal force of an acceptance condition, the buyer's intent to purchase and the rate of conversion into a full contract are none of them guaranteed by the primary sources; what this article presents is A&A's design proposal.

Anthropic

From the sources

Lovable's page states that the company reached 200 million dollars in annualized recurring revenue twelve months after launch and is now at 400 million dollars ARR, that more than 50 million projects have been built on the platform at over 200,000 per day, and that apps built there draw more than 600 million visits a month. Uber, HubSpot and Zendesk are named among the large enterprises that build on it.

Anthropic

A&A perspective

These are figures the vendor published itself, without the population, period, definitions or comparison conditions being shown, and they are not independent verification. They cannot be used as a forecast for what a one-person or few-person company in Japan would achieve taking on the same work, so this article extracts only the thinking about evaluation design, not the numbers. There is another difference that matters. Lovable is a product holding tens of millions of projects and can absorb what it gets wrong on hard cases across the whole. A single customer's paid pilot has no such buffer. Miss three of ten and it is received not as thirty percent but as "it did not work". Both sources are written about evaluating products, and reading them across into clauses of a service contract is A&A's hypothesis.

AnthropicAnthropic

A&A perspective

The next step. If what to sell as the first product is not settled yet, read "Choosing your first paid AI service: how a solo founder carves out one sellable unit of work" first. How to build the completion check itself as an implementation is separated into "How to verify the business result after AI says it is done". Where this kind of test sits inside the whole flow from acquiring customers to keeping them is set out in "AI-native GTM: a practical guide for solo founders and small teams". If the stage you are at is organising intake from inquiry to quotation, "Wire inquiry-to-quote with AI: fix missing information first at a small service firm" is the closer match. Once the target work is settled on one thing and the ten rows can be filled, you can scope development on exactly that range.

Write only features, duration and price into a paid-test proposal and nobody is in a position to declare pass or fail when it ends. Let the buyer pick the cases, list up front the cases the AI should hand to a person as a kind of pass, and finish exactly one case yourself before pricing. Name one person by role to check, and put the hours they will spend into the line items. Add who bears re-checking when the model or procedure changes, and ten rows are enough. But at a stage where the buyer has no budget yet, these ten rows are premature: hand over a short unpaid demo and one page of scope first.

Sources & editorial note

Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.

  1. Lovable Claude Platform (API) case study

    Anthropic · n.d. (no date shown on page)

    Accessed 2026-09-23
  2. Demystifying evals for AI agents

    Anthropic · 2026-01-09

    Accessed 2026-09-23

AI-assisted editorial production

A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.

Editorial check: 2026-09-23

All articles