Skip to main content
Navigation

A&A INSIGHTS

Business & AI strategyFor business owners

No Case Studies Yet: Your Operating Record and a Check Set

What evidence a proposal carries before you have case studies: your own operating record, or a check set built from real failures, and where to draw the disclosure line.

AI contract workproposalsno case studiesacceptance criteriafounder sales
日本語で読む
Scattered documents become an organized comparison and a decision
A conceptual illustration of gathering information, organizing it, comparing conditions and making a decision. Illustration generated with AI

THE STARTING POINT

When you have no case studies yet, what belongs in a proposal is not an outcome number but material two people would independently judge the same way. There are two kinds: a record of work you actually ran inside your own business, and a check set built from failures you actually hit, with explicit pass and fail criteria. Both can be handed over without naming a customer, and the buyer can reach the verdict on the spot. A discount is neither, so it does not fill the gap.

Instead of a case study, show evidence two people would judge the same way

A&A perspective

When you have no case studies yet, what belongs in a proposal is not an outcome number but material two people would independently judge the same way. There are two kinds. A record of work you actually ran inside your own business, and a check set built from failures you actually hit, with explicit pass and fail criteria. Both can be handed over without naming a single customer, and the buyer can reach the verdict on the spot. A discount is neither of these, so it does not fill the gap a missing case study leaves.

From the sources

Anthropic's guidance on evaluations for AI agents, published 09 January 2026, defines a good task as one where "two domain experts would independently reach the same pass/fail verdict". In the same section it states that "Ambiguity in task specifications becomes noise in metrics".

Anthropic ↗

A&A perspective

Both sentences were written about measuring the quality of AI agents, not about assembling sales material. A&A reads them across as a test for admitting evidence into a proposal. Read that way, a case study with a customer name does not pass the test: the buyer cannot re-run it and therefore cannot reach the verdict themselves. What is actually missing at the no-case-study stage is not the case study. It is material shaped so the other side can judge it without you.

What the buyer is really checking when they ask for case studies

A&A perspective

When that question comes up in a meeting, the buyer is not asking for another company's name. They are usually checking three things: whether the thing they are ordering actually runs, who fixes what when it does not, and who decides what counts as done. A case study answers those three indirectly. It is not the direct answer. That is why material answering the three directly can stand in for a case study you do not have.

A&A perspective

The reverse also holds: a proposal that carries case studies but never answers the three stays weak. "Deployed at a major firm" tells the buyer nothing about what was called a pass in that engagement, so they are left guessing the pass criteria for their own. Trying to compensate for missing case studies with company names and scale moves you further from the three questions. What you should be filling in is whichever of the three you can write down today.

Proposal material, re-sorted by the buyer's question and the condition that makes it evidence
Material in the proposalThe buyer's question it answersWhat makes it hold as evidence
A case study with a customer nameIs it used at other companies?The buyer cannot re-run it, so they cannot reach the verdict. It also needs disclosure consent, and at this stage you have none to give.
A record of work you ran yourselfDoes it actually run?The work you ran sits close to the buyer's, and the parts of the tooling and procedure the buyer could reproduce are stated explicitly.
A check set built from failuresWhat will be called a pass?Pass and fail fit in one sentence, two people would reach the same verdict, and both the items to pass and the items to stop are included.
A discountNone of the questionsDoes not hold. The ambiguity stays, only the price moves, and a second doubt appears about why it is cheap.
Generic technical description or a list of model namesNone of the questionsDoes not hold. The same wording fits any company's proposal, so there is nothing to reach a verdict on.

Evidence A: the record of work you actually ran yourself

From the sources

On 17 February 2026, Gumloop published an account, under Max Brodeur-Urbas, of building its own customer support operation on its own platform, using the same tools open to any Gumloop user. The piece opens by saying the team "built a fully automated support operations system, using the same tools available to every Gumloop user", and closes by repeating that "Everything described above was built using the same Gumloop platform available to every customer". On staffing it states that "our support team has only two people".

Gumloop ↗

From the sources

The same piece is explicit that the system did not arrive complete: "it grew one workflow at a time, over months of continuous iteration". The units it describes are concrete. The agent watching platform health, for instance, escalates under a stated condition — "If (and only if) a finding is actionable, it alerts a human". In its closing paragraph it writes that "The support team saw problems and were able to build their own solutions".

Gumloop ↗

A&A perspective

This is not a customer case study; it is a company publishing its own operation. What A&A takes as transferable is not the presentation but the condition that makes it readable as evidence. It works because the piece states the system was built with the same tools available to every user, so the buyer can reach the same components. Inverted, this tells you what to write when you put your own operating record into a proposal: name the parts of the tooling and the procedure the buyer could reproduce. Anything resting on internal data or infrastructure only you hold does not function as evidence, because the reader cannot map it onto their own setting.

Evidence B: a check set built from failures you actually hit

A five-row, two-column comparison table headed "Which one leads depends on distance". The columns are Evidence A, your own operating record, and Evidence B, a check set built from failures. Row one, "Buyer question it answers": Evidence A answers whether it actually runs; Evidence B answers what counts as a pass. Row two, "What makes it hold": Evidence A holds when your work is close to the buyer's; Evidence B holds when two people agree on the verdict. In other words, Evidence A becomes evidence only once the parts of the tooling and procedure the buyer could reproduce are shown, and Evidence B becomes evidence only when pass and fail are written so the verdict does not split. Row three, "What you build it from": Evidence A from a setup you kept running; Evidence B from manual checks and real failures. Row four, "What you do not show": Evidence A, the real input data; Evidence B, anything naming a counterparty. Row five, "When it stops working": Evidence A, when the work you sell is distant from the work you ran; Evidence B, when you count only passes. The point of the figure is that the two kinds of evidence answer different buyer questions, so which one leads is decided by the distance between the work you are selling into and the work you actually ran. Gumloop stating that it built its support operation with the same tools available to every user, that the system grew one workflow at a time over months of continuous iteration, and that what it publishes is the components and the division of roles rather than customer names or ticket contents, all come from the Gumloop article. Beginning an eval suite from the manual checks already run and from real failures, and the point that one-sided evals create one-sided optimization, come from the Anthropic article. Evidence B's condition of verdict agreement follows the source; Evidence A's condition of closeness, the rows for what you build it from, what you do not show and when it stops working, and the framing of these two as proposal evidence are A&A's design proposal. This figure shows no effect measured by A&A.

From the sources

The same Anthropic guidance puts the starting point for an eval suite in work that already exists rather than in new authoring: "Begin with the manual checks you run during development", and, if you are already running in production, "look at your bug tracker and support queue". On size it says "20-50 simple tasks drawn from real failures is a great start", explaining that early on each change has a clear effect and that "this large effect size means small sample sizes suffice".

Anthropic ↗

From the sources

It also names the cost of delay: "Evals get harder to build the longer you wait". The reason given is that early on product requirements translate naturally into test cases, while waiting leaves you reverse-engineering success criteria from a live system. It further notes that it is useful to prepare, for each task, "a known working output that passes all graders" — a reference solution that proves the task is solvable and that the grading is wired correctly.

Anthropic ↗

A&A perspective

A&A reads this across as follows. You build a check set to control the quality of what you deliver, but the moment it exists it is also material you can put in a proposal. A single page saying "we will verify these twenty items, against this data, at this threshold, before delivery" lets the buyer reach a verdict on the spot without a single customer name appearing on it. It also cannot be copied out of a competitor's proposal, because it was derived from failures you personally hit. One caution: the source's twenty-to-fifty figure is about eval suites for agents, not a standard for how many items a proposal should list.

Admit evidence on one test: would two people reach the same verdict?

From the sources

The source is specific about grader design too. On testing only one direction it states that "One-sided evals create one-sided optimization", illustrating this with search: testing only the queries where the model should search can yield a model that searches for almost everything. On constraining the route it states that it is often "better to grade what the agent produced, not the path it took".

Anthropic ↗

A&A perspective

Put each piece of proposal material through the test one at a time. The procedure is to set that piece alone on the desk and ask yourself whether a different person looking at it would arrive at the same pass or fail. "High accuracy" does not survive, because nothing says what counts as correct. "Across 100 invoices, count how many have amount, invoice date and counterparty name all matching the source document; 95 or more is a pass" does survive, because the judge, the object, the counting rule and the boundary number are all present.

A&A perspective

The source's point about one-sided optimization carries straight over to a proposal's check set. A check set that counts only the items processed correctly will wave through a build that also processes the items it should have refused. If you are putting a check set in front of a buyer, include both what should pass and what should be stopped. A buyer who knows the work reads your experience off whether the stop side is there at all. This is not about how the evidence looks; it is about which standard you will be held to after you win.

Anthropic ↗

Decide what you will and will not show before the meeting

From the sources

What Gumloop's account publishes is the set of components and the division of roles — automatic enrichment of user information, error detection and context attachment, issue triage, response drafting, and follow-up. It goes further, naming the agents ("Gummie Support", "Support Captain"), the third-party services wired in, and which model is used for which kind of task. The information attached is the user's profile, activity history and error counts, refreshed every 30 minutes. What does not appear: customer names, the content of individual tickets, and any pass/fail standard for the output of any unit.

Gumloop ↗

A&A perspective

A&A's recommendation starts with a caution: what Gumloop shows and withholds is not a template to copy, because that article states no pass/fail standard for the output of any unit. What goes into a proposal is the components plus the pass criteria. With that said, fix the line before the meeting rather than during it. Show the units of the system and the pass criteria for each unit. Do not show the real data used as input, any description that identifies a counterparty, or any part you have not yet reproduced yourself. The practical gain from deciding this in advance is that when someone presses for "a bit more detail", you are not choosing on the spot between going quiet and saying too much. For a founder running the meeting alone, removing one in-the-moment judgement is not a small thing.

A&A perspective

This article calls the one-page version of all of this an evidence card. It has five fields: (1) the buyer's question, (2) the evidence you offer, (3) the judge and the threshold, (4) what you show, (5) what you do not show. Prepare one blank before the meeting; any field you cannot fill is a gap in your preparation. In particular, while field 3 is empty, that material is not ready to go into a proposal. The hypothetical below shows the card filled in.

Why a discount does not fill the gap

A&A perspective

Discounting because you have no case studies does not answer the buyer's question. What the buyer is carrying is uncertainty about what will be called done, and a lower price adds not one word to that definition. If anything, the drop invites a second doubt on their side: that it is cheap because something is missing. You have removed none of the original uncertainty and added a new one.

A&A perspective

There is a further, concrete loss. Win the work without written criteria and there is nothing to settle acceptance against, so revisions continue until the other side feels satisfied. A reduced unit price combined with an unbounded number of revisions is the worst shape an early engagement can take. Leading with a check set is not a display of good faith; it is the practical move that closes the scope of rework in advance. In order: the pass criteria come out before the price conversation.

Hypothetical: the evidence card of a solo developer bidding on invoice reconciliation

Hypothetical example

What follows is a hypothetical A&A constructed for explanation, not an actual engagement. The setup: the supplier is a single developer whose past deliveries cannot be disclosed. The prospect is a building-materials wholesaler that reconciles supplier invoices against its own purchase orders every month, with two staff handling roughly 400 invoices a month. The prospect opened the meeting by asking for case studies from the same industry.

Hypothetical example

Filling in the five fields of an evidence card. (1) The buyer's question: can amount discrepancies in incoming invoices be caught before a person checks them? (2) The evidence offered: the same configuration the developer has been running for three months on their own expense reconciliation — extract amount, date and counterparty name from the PDF, compare against the ledger, and list only the differences — plus a check set built from the 28 cases it actually got wrong over those three months. (3) The judge and the threshold: the prospect's own staff feed in 100 of their invoices and count both how many have all three fields matching the source document and how many discrepancies were missed; 95 or more matching with zero missed is a pass. (4) What is shown: the list of extracted fields, the comparison procedure, the wording of all 28 check items, and the counting rule. (5) What is not shown: the developer's own expense data, counterparty names, and the parts of the internal correction rules the buyer could not reproduce.

Hypothetical example

What does the work in this hypothetical is that 28 is a count of failures. Those 28 items cannot be derived from processing that went well. Without having looked at the 28 cases it got wrong, the check set would list only items anyone could write, such as "the amount matches". Being able to volunteer the zero-missed condition comes from the same place: knowing where the misses happen. Note that every number here — the counts, the three-month period, the pass threshold — is a setting chosen for explanation. None of it is a value A&A measured, and none of it is a commitment to reach that level.

What does not transfer

From the sources

Gumloop's article presents its scale figures as its own operation: over 500,000 support-related workflows a week, 18 unique MCP tools, and response times under five minutes. What the piece sets out is the configuration of that system and the history of how it accumulated one workflow at a time.

Gumloop ↗

A&A perspective

The limits from here are A&A's. Those figures describe that company's own operation and are not grounds for another company to project the same. The piece was published in the context of explaining its own product, and it is not an independent audit. Nor is it a measurement that the same configuration produces the same result elsewhere. The article itself states none of these caveats, so they are qualifications A&A is adding.

A&A perspective

A record of your own operation works as evidence only when the work you are selling into sits close to the work you actually ran. A record of running expense reconciliation carries weight in a proposal about invoice reconciliation; in a proposal about handling inbound support it is only evidence that you can operate the tooling. Force a distant record into a proposal and the buyer notices the distance rather than the content. In that case, lead with Evidence B and write down what counts as a pass in their work first.

A&A perspective

What this article does not establish should be stated plainly. A&A holds no measurement of how presenting a check set moves win rates or unit prices. Nor has A&A observed whether, in Japanese deals where buyers ask for case studies, presenting a check set actually helps hold price. Both sources are documents overseas companies published about their own technical work; neither was written about selling contract development in Japan. What this article proposes is only an order of consideration: admit evidence on whether two people would reach the same verdict.

What to do next

A&A perspective

Pick one. If something has already been running in your own business for several months, lead with Evidence A and write it out down to the components the buyer could reproduce. If nothing is running yet, start from Evidence B. The starting point the source recommends is not new test items but the checks you already run by hand and the failures that actually occurred, so you can begin by collecting the cases you recently got wrong. The longer you leave it, the more the work becomes reverse-engineering criteria out of a running system. Either way, fill in one evidence card, all five fields, before your next meeting.

Anthropic ↗

A&A perspective

If you have both and cannot decide which to lead with, the confidentiality line is usually what settles it. When what you may show is narrowed by your sector or by past contracts, working through those constraints once is often the faster route. A&A's initial consultation is free and can be used purely to settle the approach. If the thing you want built is already decided, you can go straight to a development quote without the advisory step.

What is missing at the no-case-study stage is not a case study; it is material the other side can judge for themselves. Two kinds of evidence survive that test: a record of work you actually ran in your own business, and a check set built from real failures with explicit pass and fail criteria. Which one leads depends on how close your own work sits to the buyer's and on how much you are free to show. A discount is neither, so what goes out first is not a price but a pass criterion.

Sources & editorial note

Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.

  1. Supporting the world's most AI-native companies with a 2-person team

    Gumloop · 2026-02-17

    Accessed 2026-10-04
  2. Demystifying evals for AI agents

    Anthropic · 2026-01-09

    Accessed 2026-10-04

AI-assisted editorial production

A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.

Editorial check: 2026-10-04

← All articles