A&A INSIGHTS
Outcome pricing for AI agents: how do you count one completion?
Outcome pricing starts with a definition of one completion the customer can reconcile, not the unit price. Decagon and Stripe's metering docs build the six-line sheet.
日本語で読む
THE STARTING POINT
One completion counts only when something the customer can verify in their own records has changed. Handoffs to a person and out-of-scope requests are shown on their own lines rather than billed as completions, repeat attempts at the same request collapse into one, and a completion arriving after the finalization grace period counts in the following month.
Before deciding how to count, decide what makes one completion
A&A perspective
The first decision in outcome pricing is not the unit price but the definition of one completion. A case counts as one only when something has changed that the customer can look at in their own records and confirm as finished. Handoffs to a person and requests judged out of scope do not count as one; they are not erased but counted on their own lines of the statement. A completion that arrives after the finalization grace period is not excluded either: it is counted as one completion in the following month. Repeat attempts at the same request collapse into a single completion, which is what prevents double billing. A reversal after completion leaves the completion counted and is returned on its own line, and every case carries a reconciliation key so the two records can be matched. That is the answer this article gives; what follows is the evidence for it and the six lines to fill in before the contract. If you bill the number of AI replies as though it were the number of resolutions, then the moment retries and human handoffs mix in you can no longer explain what you are charging for. That the first design decision in outcome pricing is a reconcilable definition of completion rather than a price is an A&A editorial position, not a statement about any particular price or contract term.
From the sources
Anthropic's published case study on Decagon states the distinction plainly. Quoting the company's CTO and co-founder, it says that it is not just about answering questions and that their AI agents actually complete tasks for customers. This is a vendor account published by Anthropic on its own site, not an independent third-party audit.
The Decagon case: stopping at instructions, or finishing the work
From the sources
The same quote continues into a concrete illustration. Rather than giving instructions for a refund, Sreenivas says, "the agent checks your eligibility, processes the refund through the payment system, and handles all the steps a human agent would have done." From the customer's perspective, he adds, it is as good as talking to a skilled support representative.
From the sources
The same page also states that with Claude, end-users resolve issues without human intervention "in most cases". The top of the page additionally displays a 70% reduction in over-inferencing rates, a vendor-published internal metric about over-inferencing rates rather than a count of completions.
A&A perspective
Placing those two side by side draws the line outcome pricing needs. A completion is something that changed in the customer's world. In the refund example, completion is the state where the refund has been executed and appears in the customer's transaction record. The phrase in most cases, meanwhile, says there is always a remainder. Where human review is the expensive step, cases escalated to a person carry heavy hours on your side while delivering none of the promised outcome. A&A's proposal is not to leave that remainder unbilled, but to count it from the start as a separate unit. Hide the handoff count and, in A&A's reading, your pricing gets worse the more human work a case takes.
| What happened | Billed as one completion? | Why |
|---|---|---|
| Processing finished and the result appears in the customer's own records | Yes | The customer can confirm the completion from their own side |
| A reply was returned but the customer asked for a human | No. Count it in a separate unit | The promised outcome was not delivered |
| The same request arrived twice and was processed twice | Count it once | An idempotency key removes the duplicate so the customer is not billed twice |
| The customer reversed it after completion | Count it, and show the reversal on its own line | The work was actually performed |
| The completion record arrived after the finalization grace period | Count it in the following month | Usage reported beyond the grace period is not included on that invoice |
| There is output claiming completion but no execution record | No | A statement is not evidence of execution |
The usage meter is not the record of truth for completion

From the sources
It is worth looking at the billing side too. Stripe's documentation on recording usage states that Stripe processes meter events asynchronously, so aggregated usage in meter event summaries and on upcoming invoices might not immediately reflect recently received meter events. It also says you can decide how often you record usage, whether as it occurs or in batches.
From the sources
The same documentation instructs you to use idempotency keys to prevent reporting usage for each event more than one time because of latency or other issues. It also requires that the timestamp is within the past 35 calendar days and not more than 5 minutes in the future, explaining that the five-minute window covers clock drift between your server and Stripe's systems. Timestamps outside that range appear among the listed meter error codes as timestamp_too_far_in_past and timestamp_in_future.
A&A perspective
What follows is that the meter total is an aggregation for billing, not the record of truth that work finished. Because it is assembled asynchronously, the sum at any given instant is not final. The very existence of idempotency keys shows that duplication is treated as expected rather than exceptional. A&A's reading is that the record of truth should live in your own operational log, flowing one way into the billing meter. If you use the billing figure as your operational record instead, you cannot trace what actually happened when a customer questions an invoice. And unless the unit you send matches the completion defined at the start of this article, no reconciliation is possible at all.
Which month does a late record belong to?
From the sources
The close has its own specification. Stripe's documentation says that all invoices have a default finalization grace period of 1 hour, and that during that grace period you can continue to report usage for the previous billing period. It also narrows where that matters: only subscription cycle invoices, and subscription schedule phase transition invoices that modify metered items, include usage reported during the grace period. Every other invoice reflects only the usage accrued up to the point the invoice was created. During finalization, Stripe updates the invoice to reflect the latest quantity for its billing period.
From the sources
The same documentation states that any usage reported beyond the grace period is not included. For invoices enabled for automatic collection the grace period can be extended up to 72 hours, but the documentation warns against specifying a value longer than your service period, noting that a daily service period should not carry a grace period of 24 hours or more. For an invoice with several products that might satisfy more than one rule, the more conservative grace period applies.
A&A perspective
For a service where work runs overnight, or where a human review means a case completes the following morning, the length of that grace period moves the month-end count directly. A&A's proposal is to stop treating it as a technical setting, align it with the cut-off time stated in your agreement, and state it in the customer-facing explanation too. Completions that arrive after the grace period should then be counted in the following month rather than quietly pushed into the current one. Pushing works once; the moment the customer reconciles against their own records, it surfaces as a discrepancy, and the credibility of the whole invoice goes with it. Discrepancies themselves are unavoidable. What is avoidable is being unable to explain them.
The completion-definition sheet: six lines
A&A perspective
The argument collapses into six lines to fill in before the contract. First, what marks the start of a case: the arrival of the customer's request, or the point at which you begin processing. Second, what has to change on the customer's side for it to count as complete, expressed as a change they can verify in their own records. Third, what is excluded from completion: handoffs to a person, requests judged out of scope, and output with no execution record behind it. A completion that arrives after the grace period does not belong here; it is not excluded but moved to the following month. Fourth, the treatment of retries: not counting the second and later attempts at the same request, and which key removes duplicates. Fifth, the treatment of reversals: whether a post-completion reversal is counted and shown back on its own line, or removed. Sixth, the reconciliation key, the identifier used to match against the customer's records.
Hypothetical example
Here is the sheet filled in with an invented case. Suppose the service reconciles incoming payments against issued invoices on the customer's behalf. The start is the moment an unreconciled payment appears in the customer's accounting system. Completion is the state where the match is confirmed and the customer's outstanding balance has fallen by that amount. Excluded from completion are payments routed to a person because the match was unclear, and those held because the payer name did not agree. For retries, second and later attempts against the same payment ID are not counted, and the payment ID serves as the idempotency key. For reversals, if a match is later undone, the completion stays counted for the month and a reversal line is added separately. The reconciliation key is the customer's own payment ID. This is an invented setup that shows the shape of the design; it is not a real customer record and not an outcome delivered by A&A.
Make a statement the customer can recompute
A&A perspective
Once the definition exists, decide the shape of the statement attached to the invoice. What is needed is not a single total but the material from which the customer can reproduce the same number out of their own records. A&A suggests five figures side by side: completions billed for this month, cases handed to a person, cases judged not billable (requests out of scope, and output with no execution record), cases that arrived after this month's grace period closed and carry into next month, and cases carried in from the previous month that are billed here. With those five present, any gap between the total the customer counted and the number on the invoice can be explained inside the statement itself. Breaking out the counts looks like an invitation to negotiate the price down; A&A's expectation is the opposite, because it is the single unexplainable number that becomes the negotiating material.
Hypothetical example
To continue the invented reconciliation service: the customer counts 420 incoming payments for the month in their own accounting system. Of those, 19 were handed to a person, 9 were judged not billable, and 4 completed after the grace period closed and carry into next month, leaving 388 billable completions for the month. Three completions carried in from the previous month bring the invoice to 391. With the five figures on the statement the customer can close the check: 391, minus the 3 carried in, plus 19 handoffs, 9 not billable and 4 carried out, returns the 420 they counted. Had the statement carried only the number 391, their staff could not have explained the 29-case difference internally, and verification of the invoice would simply stall. These figures were constructed for the explanation; they do not represent real volumes or any measured accuracy.
What this article does not decide
A&A perspective
Three limits. First, a specification decides none of the commercial questions. The Stripe documentation quoted here explains how usage is recorded and when a cycle closes; it says nothing about what one completed outcome is worth or what a fair unit price would be, and neither does this article. Nor is the metering implementation itself the subject here; the question is only what counts as one unit sold. Second, the Decagon case is a vendor-published account rather than an independent audit. Its figures are Decagon's own published metrics for its own product and operating conditions — the page lists the company size as Small — and are not a forecast for another team. The source's own wording is "in most cases", and the residual human path always carries cost. Third, the reconciliation format proposed here is an A&A proposal, not an operating procedure observed anywhere. All quoted specifications reflect the documentation as read and should be reread before implementation.
A&A perspective
The next step is to open the invoice you send today and check whether a customer could reproduce its count from their own records. If they could not, return to the six lines before arguing about the price. If you first want to see where this billing decision sits between acquisition and continued use, "AI-native GTM: a practical guide for solo founders and small teams" provides that map. If the nearer problem is not the unit price but how to show usable volume and where extra charges begin, "How to explain credits in an AI SaaS plan: show the usable amount and the extra spend" covers adjacent ground. When what counts as complete keeps shifting per customer and you cannot settle it internally, an advisory conversation fits. When the definition is already firm and the open problem is building the measurement and reconciliation path from your operational log to the invoice, scoping it as a development engagement finishes faster.
The first task in outcome pricing is not setting a price. It is deciding what has to change on the customer's side for a case to count as one, and sharing that definition with them. With the definition in place, the argument about unit price is short. Without it, the count on your invoice is only the number of rows in your log, and the customer has no way to verify it.
Sources & editorial note
Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.
- Decagon delivers white-glove customer service at scale with Claude
Anthropic · n.d.
Accessed 2026-09-25 - Record usage for billing with the API | Stripe Documentation
Stripe · n.d.
Accessed 2026-09-25 - Configure an invoice finalization grace period | Stripe Documentation
Stripe · n.d.
Accessed 2026-09-25
AI-assisted editorial production
A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.
Editorial check: 2026-09-25