A&A INSIGHTS
Voice AI services: separate conversation quality from task completion
Choose your first voice AI use case by whether a record survives outside the call. Where none does, all you can promise is conversation quality.
日本語で読む
THE STARTING POINT
Choose the first voice AI use case you sell by whether a record survives outside the call. Where none does, as in practice or listening support, all you can promise is conversation quality; where one does, such as bookings or intake, natural speech is only a precondition and the record itself becomes what you promise.
Choose the first use case by what record survives the call
A&A perspective
When you turn voice AI into a product, choose the first use case you sell by asking what record survives outside the conversation once the call ends. Where nothing survives, all you can promise is the quality of the conversation. Where something does survive, such as a row in a booking ledger or an intake record, natural speech stops being the product and becomes a precondition, and what you promise becomes the record itself. Until that one line is settled, neither the endpoint of your demo nor the unit of your estimate is settled. This is A&A's judgement.
A&A perspective
Why does this distinction become necessary specifically in voice? In a text exchange the content of each reply stays on screen, so the buyer can read back and see whether the request was actually finished. In voice, the signals observable during a sales meeting are limited to roughly three: was it heard correctly, did it speak naturally, were the pauses not awkward. All three are properties of the conversation. Inside the buyer's head those three passing leads straight to the conclusion that intake work can be handed over too, yet nobody opened the booking ledger during the call. If the seller leaves that leap alone, the mismatch surfaces for the first time at acceptance.
From the sources
As material for the comparison we use two customer case-study pages Anthropic publishes on its own site: the page for Hume AI, which provides a voice conversation platform, and the page for Decagon, which provides a customer support platform. Both show Industry: Software in the industry field and Company size: Small in the company size field. They are cases written by the same publisher about companies in the same size bracket.
Two cases from the same publisher lead with different kinds of numbers
From the sources
Hume's page leads with three metrics. Over 2 million minutes of AI voice conversations completed; 36% of users choose Claude, higher than any external LLM; and an 80% reduction in costs with a 10% decrease in latency through prompt caching. What is being counted is the total volume of conversation, the model selection rate, and cost and latency.
From the sources
The same page also explains Hume's reason for choosing the model in terms of conversational properties. After evaluating multiple models, the page says the decision came down to Claude's vibe and natural conversational abilities, and it quotes founder Alan Cowen saying that Claude is very eloquent and that it has a personality people enjoy talking to. The company's platform, EVI, is described as detecting nuances in a user's tone of voice and generating responses with an appropriate emotional resonance.
From the sources
Decagon's page is written inside a different frame. It leads with CTO Ashwin Sreenivas saying that it is not just about answering questions, that their AI agents actually complete tasks for customers. Among the deployments, the page says that Rippling, a back-office platform, replaced their automated decision-tree logic with a Claude-powered agent, and describes that agent's role as guiding clients through complex HR and payroll processes in line with local regulations.
A&A perspective
Placing the two side by side shows that the kind of metric follows from the choice of use case. Hume's page carries no metric counting what was created in the customer's own business records as a result of a call. That is not a defect. Most of the use cases the page lists are the kind of work where nothing survives outside the conversation. Decagon, conversely, declares up front that the point is finishing the task rather than stopping at guidance. What the seller has to choose is which family of numbers they want to be judged by. This reading is A&A's interpretation; neither company states it.
| Use case | Record surviving outside the call | What you can promise when you first sell it |
|---|---|---|
| Practice for difficult conversations | None (what remains is the participant's own sense of it) | Conversation quality, and how many practice runs were completed |
| Listening and consultation desk | None (whether a summary is kept is a design choice) | Conversation quality, and where the summary is delivered |
| AI tutoring | None by default (a record exists only if you build a learning log) | Conversation quality, and the decision whether to build a log |
| First-line response to enquiries | Contact history, and whether it was handed to a person | That the matter and its history are complete at handover |
| Booking intake and changes | A row in the booking ledger | That the correct row lands and can be checked afterwards |
| Reception and identity confirmation | The intake record and the result of verification | That a record remains and the basis of verification is traceable |
A long call flips between success and failure depending on the use case
From the sources
Hume's page places conversation length on the side of results. Users have conducted over 1 million distinct conversations totalling nearly 2 million minutes, with an average conversation lasting 3 minutes and many extending beyond 30 minutes. Developers have created hundreds of thousands of custom EVI configurations, and within those Claude was chosen most often, accounting for 36% of all specified model selections.
From the sources
The page also goes into how to measure. It says Hume encourages its own customers to look beyond traditional metrics to measure impact through the lens of user wellbeing, and continues that they track not only customer satisfaction but also how interactions affect users' overall experience over time.
A&A perspective
This is where the trap lies when you sell voice AI as client work. A conversation running past 30 minutes is a good number in a practice or listening use case, because it shows depth of engagement. But if a 30-minute call happens at a dental clinic's booking desk, that is a failure. The same observed value inverts its sign with the use case. That is why a conversation-side metric cannot be borrowed as a business-side metric. Showing in a demo that a 30-minute conversation continued naturally is, when you are selling booking intake, a demonstration of your own weakness. This is A&A's assessment.
A reply that stops at guidance is hard to notice in voice
From the sources
Decagon's page states concretely how conventional automation fails. Quoting Sreenivas, it says most automated solutions either point you to a generic help article or give you a list of steps to follow, and that even recent AI-powered tools feel impersonal and put the burden on the customer to do all the work.
A&A perspective
In voice, this failure shape is discovered late. A reply shaped like guidance satisfies every axis on which voice is judged — intelligibility, naturalness, timing of pauses — so it loses no points in a sales demo. A reply that in text would read back as having finished nothing passes, when spoken fluently, as attentive service. The buyer notices the gap at the point where, after deployment, their own staff are calling people back about the same matter. This is A&A's interpretation.
Hypothetical example
Set them side by side in a hypothetical. Asked to change a booking, a voice AI answers fluently: a change to your appointment, of course; please log in to our clinic's member page, select the relevant date in your booking list, and make the change there. Judged on voice quality alone it is full marks. Meanwhile nothing has happened in the booking ledger. Given the same request, a reply of: I have moved it to 2pm next Tuesday and will send you a confirmation email, together with a rewritten row in the ledger, means the work is finished. Voice-side criteria cannot tell these two apart; only looking at the ledger can. This is an example A&A constructed, not a record from an actual customer.
A use-case sorting table: what survives outside the call
From the sources
Hume's page enumerates the use cases its Claude-powered EVI enables: practice sessions for difficult emotional interactions, mental health support conversations, customer service interactions, AI tutoring, and personal digital assistants. The page also gives the example of a customer using it for immersive coaching simulations in which managers practise delivering feedback to a defensive direct report.
A&A perspective
Read that list again through the presence or absence of a record. Practice, mental health support and tutoring leave no business record outside the conversation once the call ends. What remains is the participant's own sense of how it went. So Hume encouraging its customers to measure through wellbeing is consistent with the use cases that company chose. Customer service and personal assistants, by contrast, mix matters that do leave a record with matters that do not. Sell something mixed as a single use case and you cannot write acceptance conditions for it.
A&A perspective
So write your candidate use cases out in three columns: the use case, the record that survives outside the call, and what you can promise when you first sell it. Any use case whose second column comes out blank gets separated as a use case where you can promise only conversation quality. The separation is itself the decision. It does not mean refusing to sell the blank ones; it means not attaching a promise of completed work to them. The table below is an example of A&A's arrangement.
Put the demo's endpoint at the record, not at the end of the call

A&A perspective
If you have decided to sell a use case that leaves a record first, move the endpoint of your demo. Make the endpoint the moment a row lands in the record, not the moment the call ends naturally. Concretely, finish the demo by showing the record rather than the conversation. Share your screen, open the booking ledger, point at the row the call just created, and explain which fields are filled — date and time, name, assigned person, notes — and which are still empty. Empty fields are not a weakness. Showing them first lets the conversation move on to whether filling them is inside the contract's scope. This is A&A's proposal.
Hypothetical example
Write that endpoint out for a hypothetical dental clinic. The matter the demo puts through is one booking change from an existing patient. The endpoint is the state in which the clinic's booking ledger holds the changed row, the previous row remains marked cancelled, and the call recording and transcript can be traced from that row. Write down just as many things you are not promising: first-visit enquiries are out of scope, insurance card verification is out of scope, and reconciling a duplicate request from the same patient is out of scope. With these six lines, whoever watched the demo can judge what would happen at their own clinic. This is an example A&A constructed, not a record from an actual clinic.
A&A perspective
Before a contract you often cannot write into the customer's ledger. In that case, change the destination from the customer's production ledger to an empty ledger carrying the same columns. As long as the column names and types match production, you can still demonstrate the fact that a record lands. The work of connecting to production, and who checks it afterwards, belongs to the paid validation stage. That stage's design is covered in "What to promise in a paid AI pilot: separating the demo from the acceptance conditions". The procedure for assembling the demo itself as an eval set is in "Put the buyer's task in your AI demo: build a comparable eval set".
Price voice quality and latency as preconditions, not as the product
From the sources
Hume's page also gives numbers for technical improvement. It leads with an 80% reduction in costs and 10% decrease in latency through prompt caching, and the body text puts the latency figure at 10% or more. The same page also gives the reason it expects voice to become the primary interface between people and AI as voice's inherent advantages in speed and emotional expression.
A&A perspective
Latency and voice quality are not the product in a use case where you promise completed work. They do, however, operate as preconditions. If the reply is late the other party starts talking, the transcription of the overlapped audio breaks down, the appointment date is not captured, and no row lands in the record. So latency is a conversation-side property that nevertheless degrades a business-side result. In terms of ordering an estimate: put the landing of the record in the scope of the promise first, then write the required standards for latency and intelligibility on a separate line as the conditions that make that promise hold. This is A&A's proposal.
A&A perspective
This split has a side effect. When a request arrives to make the voice more natural, you become able to ask back whether that raises the accuracy of the record or improves how it feels to listen to. The former is inside the promise; the latter is a separate request. The design of the stage that decides the unit of completion itself is covered in "Outcome pricing for AI agents: how do you count one completion?".
Where this split does not fit, and the next step
From the sources
State the nature of the sources plainly. Both pages read here are customer case studies Anthropic published as adoption examples for its own products, not verification by an independent third party. Every figure presented comes from the publisher and the company concerned. Neither page displays a publication date, so we treat them as the text as read on 30 September 2026. The deployments Decagon's page lists include answering writers' questions on the publishing platform Substack and guiding HR and payroll processes at Rippling, and it also names Eventbrite among its customers. None of them is written as a voice-response case.
A&A perspective
Three things are therefore A&A's adaptation: carrying the distinction between stopping at guidance and finishing the processing over into voice; sorting Hume's list of use cases by the presence of a record; and reading that sorting across into line items on a client-work estimate. Hume is also a provider of a voice platform for developers, not a firm delivering voice response to Japanese small and medium businesses as client work. The list of use cases on its page consists of configurations that company's customers built, not a distribution of demand in Japanese B2B client work. Nor is any claim made that measuring by conversation length or wellbeing is inferior. For the use cases that company chose, it is a consistent way to measure.
A&A perspective
There are situations where this split does not fit. If your first counterpart is at the stage of just wanting someone to answer the phone, raising the design of records first will stall the conversation. In that case agree with the record column left blank, and write in the estimate that it is blank. Also, selling a use case that leaves a record and one that does not at the same time doubles the acceptance discussion. Narrowing to one first finishes faster. The next step is to write your candidate use cases out in three columns and separate the rows whose second column is blank. The overall ordering from acquiring customers through to retention is collected in "AI-native GTM: a practical guide for solo founders and small teams".
In a voice AI sales meeting the only observable signals are conversation-side ones. So choose the first use case you sell by whether a record survives outside the call, and if you choose one that does, put the demo's endpoint at the row in the ledger rather than at the end of the call. Natural speech then becomes a precondition rather than the product.
Sources & editorial note
Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.
- Hume AI creates emotionally intelligent voice interactions with Claude
Anthropic · n.d.
Accessed 2026-09-30 - Decagon delivers white-glove customer service at scale with Claude
Anthropic · n.d.
Accessed 2026-09-30
AI-assisted editorial production
A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.
Editorial check: 2026-09-30