A&A INSIGHTS
Put the buyer's task in your AI demo: build a comparable eval set
For small AI vendors: turn the sales demo into an evaluation set built from the buyer's anonymized tasks, run it more than once, and score pass, partial or hand back.
日本語で読む
THE STARTING POINT
A sales demo used for a purchase decision should not add more of your own success stories. Turn it into an evaluation set: three to five representative tasks the buyer has anonymized and permitted, run more than once under the same conditions, scored on three levels — pass, partial credit, and hand back to a person. A rehearsed demo is by construction the seller's own training data, and measures nothing about the buyer's work.
Make the purchase-decision demo a measurement, not a performance
A&A perspective
What you do in a sales demo that a buyer will use to decide is not add more of your own success stories. It is prepare three to five representative tasks the buyer has anonymized and permitted, run the same inputs more than once, and turn the demo into an evaluation set you can score on three levels: pass, partial credit, and hand back to a person. As long as the demo is designed as a single performance, the buyer walks away still holding the first question — what happens on my work? A polished performance only tells the buyer that you are polished.
From the sources
Anthropic's engineering article "Writing effective tools for agents — with agents" (published 11 September 2025) says, in the section describing how it evaluated its own internal tools, that its evaluations were created on top of its internal workspace, mirroring the complexity of its internal workflows, including real projects, documents and messages. It then states that it relied on held-out test sets to ensure it did not overfit to its own "training" evaluations.
A&A perspective
What A&A carries from that sentence into a sales demo is the idea of holding something out. A demo the seller has rehearsed many times is, by construction, that seller's training data: it is the result of adjusting the steps, choosing the inputs and keeping the examples that pass. Passing on the same input shows only that it passes on that input. What the buyer wants to know is the held-out side — the behaviour on its own cases, the ones that were never used for practice. If that is right, the value of a demo is not the quality of the performance but whose hands the input in the room came from.
A&A perspective
The opposing view, stated up front: a sales demo exists to reduce the other side's anxiety and keep the conversation moving, and showing something that does not work will stop the deal. There are situations where this is correct, and the last section states the conditions. The demo discussed here is the one where the buyer is already considering adoption and is looking for material to compare vendors or to explain a choice internally. At that stage, a performance that cannot be compared is weak as information.
Write representative tasks as requests that contain a judgment, not a single lookup

From the sources
The same article, in the section on generating evaluation tasks, prints strong and weak tasks side by side. Among the strong examples: "Customer ID 9182 reported that they were charged three times for a single purchase attempt. Find all relevant log entries and determine if any other customers were affected by the same issue." and "Customer Sarah Chen just submitted a cancellation request. Prepare a retention offer. Determine: (1) why they're leaving, (2) what retention offer would be most compelling, and (3) any risk factors we should be aware of before making an offer." Among the weak examples: "Search the payment logs for purchase_complete and customer_id=9182." and "Schedule a meeting with jane@acme.corp next week."
From the sources
The article says tasks should be grounded in real-world uses and based on realistic data sources and services, and recommends avoiding overly simplistic or superficial "sandbox" environments that do not stress-test the tools with sufficient complexity. It also notes that strong evaluation tasks might require multiple tool calls, potentially dozens.
A&A perspective
The difference between the two lists is not who chose the task. It is the shape of the request. The weak examples already contain the location of the answer: which log, under which condition, is written into the request itself. The strong examples describe a situation and leave it to the recipient to decide where to look. Translated to the input of a sales demo, the seller's demo script is almost always on the weak side — the order of operations is fixed, and it passes because it follows that order. The buyer's actual work sits on the strong side, because it arrives as "this message came in; work out what is going on and write what we should do next". A useful rule when writing representative tasks is therefore: write the situation and the judgment you want, and do not write the steps.
A&A perspective
The count is A&A's proposal: three to five. The figures the source gives are about a product's evaluation dataset, not about what can be brought into a single sales meeting. Below three, the buyer cannot read whether a pass was coincidence. Above five, execution and explanation do not fit in the meeting, and the conversation collapses back to one representative case anyway. More important than the count is that the three to five fail in different ways from one another. Five cases of the same kind teach the buyer as much as one.
Hypothetical example
A hypothetical example. Suppose you are proposing, to a trading company that handles industrial machinery parts, a system that drafts replies to the stock and lead-time enquiries arriving from its customers. The seller's demo script was one case: from a message that clearly states a part number and a quantity, look up the stock table and return a lead time. The three representative tasks asked of the buyer are deliberately different in kind. One where the part number is written under an obsolete name and has to be mapped to the current one. One where a single message asks about three part numbers, one of which is discontinued. One where the quantity is written as "the same as usual", so the number cannot be determined without looking at past orders. In all three, the request text says only "draft a reply to this message" and does not say where to look. This is an invented setting for explanation; the industry, part numbers and counts are all hypothetical.
| Usual sales-demo habit | The question left with the buyer | The evaluation-set replacement |
|---|---|---|
| Show three success cases the seller picked | What happens on my cases? | Use three to five representative tasks the buyer anonymized and supplied |
| Run each case once and it passes | Did it only pass by chance? | Run the same input three times and show the spread without averaging it |
| Follow the on-screen steps in a fixed order | Is it a failure if the steps differ? | Ignore the path; judge the artifact that came out |
| Talk in terms of passed or failed | Where does a near miss go? | Write verdicts on three levels: pass, partial credit, hand back |
| Only cover the cases that work | What is it bad at? | List the hand-back cases up front with pass conditions on how they return |
| Polish the appearance of the output | Am I failing it over wording alone? | State in advance that formatting, word order and politeness never fail a case |
| End the result at "it worked" | How long did it take, and how many retries? | Record time taken, retries, where a person intervened, and failure reason |
| Conclude insufficient capability when all fail | Was the request text simply ambiguous? | Suspect the request text first, add the condition and rerun on the spot |
Borrowing inputs before a contract exists: settle anonymization and permission in one line
A&A perspective
When you borrow representative tasks from a buyer, the sales stage has no contract governing the handling of data. Once a paid pilot starts, the scope can be written into an agreement, but before that, asking "please send me three real cases" is normally stopped by the other side's legal or information-management people. When it stops there, the whole conversation about representative tasks stops with it. So decide how to ask before you ask.
A&A perspective
A&A's framing is to ask for the removals and the retentions separately. What comes out is anything that identifies someone: customer company names, individual names, contact details, account or contract numbers, internal identifiers. What stays is the shape of the work: the messiness of the prose, the typos, the omissions, several requests mixed into one message, the missing attachment. And when something comes out, replace it rather than delete it. Substituting an invented company name of the same length and form changes the behaviour of the processing less than collapsing it to "Company A".
A&A perspective
There is a limit to how far that substitution can go. The source recommends grounding tasks in real-world uses and basing them on realistic data sources and services, but it says nothing at all about anonymization. Stated as A&A's reading: the more identifiers you replace, the further the input drifts from reality and the closer it gets to the oversimplified environment the source tells you to avoid. Where substitution stops being representative is not something the source settles; that line is one you have to draw yourself.
A&A perspective
A&A's line is therefore: substitute the identifiers, leave the structure and the mess untouched. The only thing you may break is who the case is about, never how it is written. In practice, explain that line once to the counterpart on the buyer's side and have them do the substitution themselves — if you never receive the original data, there is no question of what you are holding. And do not leave the permission verbal: put one line in the proposal or the meeting note. "These three cases are supplied by the buyer in anonymized form, used only to evaluate within this sales discussion, and deleted afterwards." Whether that line exists changes how fast the conversation moves inside the buyer's organization. This is A&A's design recommendation, not something the primary sources support.
Hypothetical example
In the hypothetical example above, customer company names were replaced with invented names of the same character length, individual names with invented surnames, and only the leading symbol of each part number with a different symbol. What was not replaced: the phrasing of the quantity as "the same as usual", the structure of three part numbers mixed into one message, and the quotation of the previous exchange appended at the end. Because those three were kept, the first case failed the name mapping, the second skipped the discontinued part number, and the third returned the quantity blank. An invented setting.
One pass is not a result: run the same input several times
From the sources
Anthropic's engineering article "Demystifying evals for AI agents" (published 9 January 2026; written by Mikaela Grace, Jeremy Hadfield, Rodrigo Olivares and Jiri De Jonghe) states that agent behaviour varies between runs and that a task which passed on one eval run might fail on the next. It captures this with two metrics: pass@k is the likelihood that an agent gets at least one correct solution in k attempts, and pass^k measures the probability that all k trials succeed. On pass^k the article shows the calculation that if an agent has a 75% per-trial success rate and you run three trials, the probability of passing all three is 0.75 cubed, roughly 42%, and says this metric especially matters for customer-facing agents where users expect reliable behaviour every time.
A&A perspective
In that framing, a sales demo is a single trial at k=1. If it passes, it displays "passed", but all that can be read from it is that the per-trial success rate is greater than zero. What the buyer will live with after adoption is the pass^k side: does it pass every time. If a demo lets those two run together into an agreement, the conversation after adoption becomes "but it worked in the demo". The source is describing product evaluation; moving it into a sales meeting is A&A's reading.
A&A perspective
In practice: run each representative task three times on the same input, and show all three runs. List the case that passed three out of three, the one that passed two, and the one that passed none, without rounding them into an average. That list is what the buyer can actually compare. It also compares across vendors, as long as the input and the number of runs are the same. A&A sets the count at three because five or ten does not fit the meeting's time or the API spend; the source gives no basis for three. Look at the cost first: five representative tasks at three runs is fifteen executions, and for anything slow that means doing the runs before the meeting and bringing the record. One more cost is worth stating. Running the same input three times in front of someone can read as a lack of confidence. Whether to say "there is variance, so I will show you three runs" up front depends on your relationship with that buyer.
Write the verdict on three levels: pass, partial credit, hand back to a person
From the sources
The same article says that for tasks with multiple components you should build in partial credit. The example it gives is that a support agent which correctly identifies the problem and verifies the customer but fails to process a refund is meaningfully better than one that fails immediately, and it states that this continuum of success needs to be represented in the results.
From the sources
The article also addresses the common instinct to check that agents followed very specific steps, such as a sequence of tool calls in the right order, and says the approach is too rigid and results in overly brittle tests, because agents regularly find valid approaches the eval designers did not anticipate — so it is often better to grade what the agent produced, not the path it took. The tool-design article points the same way, noting that because there might be multiple valid paths to solving tasks correctly, one should try to avoid overspecifying or overfitting to strategies.
From the sources
The tool-design article adds that you should avoid overly strict verifiers that reject correct responses due to spurious differences like formatting, punctuation, or valid alternative phrasings.
A&A perspective
Moving those three into a sales verdict gives three levels instead of two. "Pass" means the buyer can hand the output straight to the next step. "Partial credit" means how far it got is identifiable and the remaining work can be specified. "Hand back to a person" is for cases where returning the work to a human without processing it is the correct behaviour, and it counts as a pass when the return states what was missing. Written as a binary, the partial-credit cases fall into the fail column, and behaviour that is actually usable becomes the reason you lose the deal. Then add two conditions on how verdicts are written: the path does not matter, and differences in formatting, word order or level of politeness do not fail a case. The only thing that fails a case is a substantive gap that stops the buyer handing it on. Putting those two lines at the top of the evaluation set means a competing vendor gets measured on the same line.
Hypothetical example
Filling that in for the three hypothetical cases. Case one, the obsolete part number: it passes if the mapping is made; it is a hand-back pass if, unable to map it, the reply states "this part number does not appear in the current parts table"; it fails if it silently answers about a different part. Case two, three part numbers with one discontinued: it passes if it answers on two and explicitly flags the discontinued one; it is partial credit if it addresses only two of the three and drops one; it fails if it quotes a lead time for the discontinued part. Case three, "the same as usual": it passes if the quantity can be determined from past orders; it is a hand-back pass if it states that it cannot be determined and asks for the order history; it fails if it drafts a reply leaving the quantity blank. None of these verdicts move on differences in tone or honorific style. An invented setting.
The columns to record, and how to read a run where everything fails
From the sources
The tool-design article lists metrics worth collecting besides top-level accuracy: the total runtime of individual tool calls and tasks, the total number of tool calls, the total token consumption, and tool errors.
A&A perspective
A&A translates that into four columns for a sales evaluation set. Time taken per case. Number of retries. Where a person had to intervene. For failed cases, the category of the reason. Token volume carries straight over as cost per case when you are selling a usage-priced product, but in a services proposal it is usually not the buyer's concern, so add it only when you need to explain cost. The effect of writing those four is that differences invisible in the verdict column show up. Three cases all passing but one taking twelve minutes, and three cases all passing in about two minutes each, are different stories in the buyer's work.
From the sources
The evals article states that with frontier models a 0% pass rate across many trials is most often a signal of a broken task rather than an incapable agent, and a sign to double-check the task specification and the graders. It also says everything the grader checks should be clear from the task description and that agents should not fail due to ambiguous specs, explaining that ambiguity in task specifications becomes noise in metrics.
A&A perspective
Decide in advance how you will behave when every case fails in the meeting. Rather than concluding on the spot that the capability is insufficient, reread the request text first. If the request said only "draft a reply" and never conveyed that the stock table was available or what the lead-time calculation rule was, then what failed was the task. Being able to add the condition there and rerun hands the buyer a different piece of information: what it needs to be told in order to work. If it still fails after the condition is added, that is the most valuable result this meeting could produce — the buyer gets to take away a "no". The source is describing product evaluation; how to behave in a sales meeting is A&A's design proposal.
Hypothetical example
In the hypothetical example, running three cases three times each: case one passed three out of three at about two minutes; case two passed one out of three, failing twice by quoting a lead time for the discontinued part; case three returned the quantity blank in all three runs and so scored as a hand-back pass. The reason for case two's failures was the same in all three runs — the list used to judge discontinued status had never been supplied. Telling it where that list was and rerunning gave three passes out of three. That the failure reason collapsed to a single category was the most concrete information the buyer took from the meeting. All figures are hypothetical values for explanation.
Where this design does not fit, and the next step
A&A perspective
The cases where it does not fit, first. There is a stage at which the buyer has no budget yet and the only thing needed internally is the single point that this kind of thing is technically possible. Bringing representative tasks and three-level verdicts into that stage asks the other side to decide something it cannot decide yet, and the conversation stalls. What is fast at that stage is a short demo and one sheet stating the scope. An evaluation set earns its keep when the buyer is considering adoption and is looking for material to compare vendors or to explain a choice internally. It also does not fit when the work in question happens only a few times a month: collecting three representative cases means reaching back half a year, and those three no longer represent the current work.
From the sources
In its table comparing methods of understanding agents, the evals article lists as a weakness of automated evals that they can create false confidence if they do not match real usage patterns. The same article says it does not take eval scores at face value until someone digs into the details of the eval and reads some transcripts.
A&A perspective
That weakness applies to an evaluation set unchanged. Three to five cases are part of the buyer's work, not all of it, and three cases passing means those three cases passed. The most the presenting side can say is "under these conditions it behaved this way", never "this will work for your operation". The nature of the sources is worth stating too: both articles are Anthropic describing its own and its customers' practice, not independent audits. Both are about evaluating AI agents and tools, and neither says anything about sales meetings, buyers comparing vendors, or the handling of customer data. Setting representative tasks at three to five, running each three times, writing verdicts on three levels, and the one-line anonymization and permission clause are all A&A's design proposals rather than conclusions drawn from the sources. Whether showing a non-attainment raises the probability of winning the work has no support in the material read here. No close rate, no win rate against competitors and no search ranking is presented in this article, and none of this is recorded as work A&A has performed.
A&A perspective
The next step. If this evaluation set moves the buyer's decision forward and a paid pilot becomes the next stage, what you promise there is a separate decision. Case selection, acceptance conditions, who verifies and for how long, and re-checking after a model change are covered in "What to promise in a paid AI pilot: separating the demo from the acceptance conditions". The evaluation set in this article is the stage before it. If what you sell is not yet narrowed to one thing, you cannot write representative tasks either; in that case "Choosing your first paid AI service: how a solo founder carves out one sellable unit of work" comes first. The earlier stage of receiving an enquiry and filling in the missing information is covered in "Wire inquiry-to-quote with AI: fix missing information first at a small service firm". Where this demo sits in the whole path from winning customers to keeping them is in "AI-native GTM: a practical guide for solo founders and small teams". If you cannot work out how to write a representative task, or which part of your own delivery should be the representative one, start the conversation there.
A sales demo used for a purchase decision should be a measurement on held-out inputs, not a rehearsed performance. Take three to five representative tasks the buyer has anonymized and permitted, write them as a situation plus the judgment you want and never as steps, run the same input three times, and score pass, partial credit, or hand back to a person. Record time taken, retries, where a person intervened and the category of each failure, and a competing vendor gets measured on the same line. When everything fails, suspect the ambiguity of the request text first and add the condition on the spot. But before the buyer has a budget, this design is too early: there, hand over a short demo and one sheet stating the scope instead.
Sources & editorial note
Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.
- Demystifying evals for AI agents
Anthropic · 2026-01-09
Accessed 2026-09-29 - Writing effective tools for agents — with agents
Anthropic · 2025-09-11
Accessed 2026-09-29
AI-assisted editorial production
A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.
Editorial check: 2026-09-29