A&A INSIGHTS
Capability up but satisfaction flat: measure requests you can now accept
Handling more work but satisfaction is flat? Decide between development and price or scope with a ledger of declined requests, read against Genspark and Chatbase.
日本語で読む
THE STARTING POINT
Satisfaction is not a proxy for capability, because it only measures the requests you accepted and that set moves upward as capability grows. Measure whether you can now accept a request you were declining six months ago; if the kinds of acceptable requests have widened, the next move is a price increase or a wider target for the work you take, not more development.
The answer: measure the requests you used to decline, not satisfaction
A&A perspective
Satisfaction is not a proxy for capability. Satisfaction only ever measures the requests you accepted, and that set moves upward together with your capability: as the service handles more, customers bring harder work, you accept harder work, and the same number comes back from a different population. So the question that tells you whether an improvement landed changes shape. Can you now accept a request you were declining six months ago? If the kinds of requests you can accept have widened, the next move is not more development. It is a price increase or an expansion of the work you target. Only when they have not widened is the direction of the improvement itself in question. One qualifier belongs in the answer: only declines whose reason was that capability was short count in this test. Declines on missing premise information, on liability or on unit economics do not move with development at all.
A&A perspective
The input for this measurement is not survey responses. It is a record of the requests you declined. Declined requests are the least recorded data in a small company: the inquiry you never quoted, the task a human quietly did end to end, the job where you accepted one part and handed the rest back. None of them appear in revenue or in a satisfaction score, and they disappear. In a small company this is where the evidence that your capability grew is densest. The rest of this article gives the record format in five columns, splits the quarterly reading into three conclusions, and decides which comes first when the range really did widen: the price increase or the wider target.
Satisfaction stays flat because only accepted requests get graded

A&A perspective
The people who answer a satisfaction survey are the people who were able to use your service, and what comes back is a grade on the work that went through. That leaves a structural gap. The request that was too hard and got declined, the request a human absorbed by hand, the job you accepted and then rebuilt: none of them reach the questionnaire. When improved capability lets you accept harder work, the accepted set itself is replaced. This quarter and last quarter are grading different populations. You believe you are reading one scale over time, but the contents underneath it moved.
From the sources
The observation that satisfaction does not move is itself on record. Kay Zhu, co-founder and CTO of Genspark, states it for the search field, and the Anthropic customer page presents his view that “Search satisfaction in the field has hovered around 80% for a decade” and explains the reason this way: “Users adapt: as the system handles more, they ask harder questions, and the satisfaction rate stays flat.” However much better the underlying system gets, the satisfaction rate does not move, because users adapt and move to harder questions. The same page states that Genspark itself “was seeing the same pattern.”
A&A perspective
The conclusion is not that satisfaction can be ignored. It is that satisfaction can stay flat both when an improvement worked and when it did not. An instrument that returns the same value in both cases cannot separate them. Neither the decision to stop developing on flat satisfaction nor the decision to keep developing on flat satisfaction is supported by that instrument alone. If you intend to change a decision based on the reading, you need a different instrument.
| Reason class for the decline | Solved by a better model or more code? | If not, what to fix next |
|---|---|---|
| Capability is short (output quality, formats handled, volume processed) | Sometimes. A re-test in three months gives one reading, not a settled answer | If it passes, price or wider target; if not, question the improvement's direction |
| Premise information does not exist on the customer's side (no written decision rule) | No. However good it gets, the rule still does not exist | A questionnaire before handover: have the customer write what counts as a pass |
| Responsibility you cannot carry (an error maps straight onto their loss) | No. A lower error rate does not move where liability sits | A contract clause: the range that passes through human review, and the range you do not carry |
| Unit economics do not work (mostly bespoke formats and exception handling) | No. It gets cheaper, but the bespoke time remains | Price and a definition of out of scope: quote separately, or a rule for declining |
From the source: at Genspark, easy and hard questions broke in different ways
From the sources
The same case page describes concretely how Genspark's earlier design broke. Under an architecture of predefined workflow nodes, the page records that “Simple questions ran through too many steps. Hard questions hit walls the workflow didn't know how to route around.” Easy questions were pushed through unnecessary steps; hard questions hit walls the workflow had no detour for. The page then adds that “The system that had taken Genspark to millions of users could no longer go where users wanted to go.”
From the sources
The page also records that the kinds of requests moved, not only their difficulty. Among the uses customers invented that nobody at the company had planned for, it notes that “Japanese seafood industry CEOs use Genspark to analyze domestic demand and source new leads.” The list of uses a provider planned for and the list of requests actually arriving are not the same list. This is also the reason for column two of the ledger below, which keeps the request in the asker's words: the requests least likely to fit your own taxonomy are exactly the ones you lose the moment you paraphrase them.
A&A perspective
That the two ends break differently matters when you choose the direction of an improvement. If easy requests are taking unnecessary steps, the thing to fix is speed and the number of hops. If hard requests are hitting walls, the thing to fix is the range you can handle. A flat satisfaction number does not tell you which of the two is happening. To know which end you are declining at, you have to hold your declined requests separated by end. An average over both ends cancels exactly the information you need.
From the source: at Chatbase, customer demand moved the product's scope
From the sources
The second case is Chatbase, a company providing a platform that automates customer support with AI. The Anthropic customer page states the difference between where it started and where it is now: “What began as a document chat interface evolved as enterprise customers pushed for more sophisticated features.” It started as a screen for asking questions of a document and changed as enterprise customers pushed for more. The reason the page gives for the change in scope is enterprise customer demand. What the provider's own roadmap said is not stated on the page.
From the sources
The current scope is also stated concretely. The page explains that the platform integrates with business systems so agents can take real actions, giving as its example tasks “from checking order status in Stripe to processing refunds with human oversight” and saying of actions of that kind “with optional human oversight for sensitive operations”. Oversight is described there as optional, not required; an earlier passage on the same page says “with human oversight” without “optional”, so the page is not consistent about whether it is required. Returning support wording and executing an operation that touches a customer's money sit on the same platform.
From the sources
The page also records a way of handling customer reaction that does not collapse it into a single number. It presents founder Yasser Elsaid saying that conversation analysis means “The AI can summarize exactly what customers want, what parts of the product they dislike, and how their sentiment changes over time.” Wants, dislikes and the change in sentiment are listed there as three separate things.
A&A perspective
A single satisfaction number is those three added together and flattened. Hold them apart and you can say afterwards which of them moved underneath a flat line. With that in place, what deserves attention is where the evidence of a widened scope appeared. It did not appear in a satisfaction trend. It appeared in the list of what the product now accepts: from questions about a document, to checking an order's status, to executing a refund. That is a capability increment written out as kinds of acceptable requests. The record a reader should build for their own business has the same shape. Only one thing differs: in a one-person or very small company, the declined list carries more information than the accepted list, because the accepted list fits in three lines.
Record what you declined: five columns
A&A perspective
A declined-request ledger needs five columns. One, the date. Two, the request in one sentence, kept close to the asker's own words; translated into your internal vocabulary you will not find the same request again later. Three, the form of the decline: fully declined, a human did all of it, accepted only part and handed the rest back, or accepted and then rebuilt. Four, the reason class: capability is short, the premise information does not exist on the customer's side, you cannot carry the responsibility, or the unit economics do not work. Five, the re-test date. Neither source states a re-test interval; three months is this article's default, not a source-derived one. One spreadsheet is enough and no dedicated tool is required.
A&A perspective
The fourth column is the load-bearing part. Of those four reasons, only “capability is short” is solved by a better model or more code. The other three keep declining the same request however much the service improves. Missing premise information calls for a questionnaire before handover; responsibility you cannot carry calls for a contract clause; broken unit economics call for a price and a definition of what is out of scope. Read “the number of declines is not falling” without that distinction and you will chase a problem development cannot solve with development. The table below splits the next move by reason class.
Hypothetical example
Fill it in with a hypothetical example. Suppose you run, alone, a service that summarizes inbound inquiry email and routes it to the right person. Say four requests were declined in April. First: also pull the amounts out of the attached PDF quote and include them in the summary. Form of decline, fully declined; reason, capability is short. Second: also judge whether the inquiry is urgent. Form of decline, accepted only part, added on the premise that a human makes the final call; reason, the premise information does not exist on the customer's side, because no written rule for what counts as urgent existed there. Third: write the scope of liability for an incorrect summary into the contract. Form of decline, fully declined; reason, responsibility you cannot carry. Fourth: produce separate output matched to the formats of three internal departments. Form of decline, accepted and then rebuilt; reason, the unit economics do not work. If you re-test the first one in July and it passes, the range you can handle really did widen. The second, third and fourth failing to pass is not evidence that the improvement did not work. They were never development problems.
A&A perspective
Decide in advance what counts as a pass in a re-test. Not that an output was produced, but that the output is usable in the work of the person who asked. How to define that pass or fail is separated into another article, “How to verify the business result after AI says it is done.” Tick the re-test column only when that check is satisfied. Loosen it and the ledger stops being an instrument for capability and becomes a record of hope.
Reading the increment each quarter: three conclusions
A&A perspective
Three months later, re-test only the rows whose reason was “capability is short.” The conclusion splits three ways. First, kinds you used to decline now pass. The range has widened, and the next move is not more development but a price increase or a wider target for the work you take. Unless a widened range is converted into price or into a buyer, the improvement stays on the cost side of your business. Second, they do not pass. This is where you finally question the direction of the improvement; the ground for doubting it is the failed re-test, not the flat satisfaction. Third, there were few “capability” rows in the first place and the declines concentrate in the other three reasons. That is not a development problem. It is a question of which to fix: the questionnaire, the contract clause, or the price.
A&A perspective
One working figure on volume, with no measured basis behind it: for a one-person company, a dozen or so declines over three months is the point at which we would start reading. If they do not accumulate, suspect that declines are not being recorded before you conclude that declines are not happening. A polite decline does not feel like a decline even to the person making it. Make it an operating rule that the moment you reply “that is difficult within the current scope, so let us discuss it separately,” you write the line. Writing the line is the cheap part. Most of the cost sits in the quarter's re-test: actually re-running a dozen or more declined requests and deciding a pass or fail for each.
A&A perspective
Classifying the requests that did arrive is handled in a separate article, “Do not send AI service support straight to the backlog: fix, explain, build.” That one decides where an incoming inquiry should go; this one uses the requests that never came in as an instrument. Keeping the two ledgers apart stops them mixing: what arrived sets your response priority, and what did not arrive decides whether to keep developing.
When the range did widen: price increase or wider target first?
A&A perspective
When the range has widened, do not raise the price and widen the target work at the same time. One test is enough: are the newly passing kinds of request work that your current customers actually send? If they are your current customers' requests, the price increase comes first, because value already reaching them carries no payment. If they are requests your current customers do not send, widening the target comes first, because the value grew but nobody is there to receive it. Reverse the order and your existing customers get a price increase you cannot explain, while new buyers get a pitch with no track record behind it.
A&A perspective
What you put in front of a customer as the ground for a price increase is neither satisfaction nor a count of improvements. It is the difference in range. Name the kinds of request you used to decline as out of scope, and show that they now pass inside the contract. For the customer this is an exchange: they pay more, and the range of their own requests that goes through widens. Ground it in your effort or your release count instead and that exchange does not hold together. Columns two, three and five of the ledger — the request, how you declined it, and the re-test that passed — are, as they stand, the material for that explanation.
A&A perspective
Where this decision sits between winning customers and keeping them is laid out on the overview side, in “AI-native GTM: a practical guide for solo founders and small teams.” This article handles one point inside it, on the border between retention and price. Where you collect the return on an improvement is better thought through separately from the moves on the acquisition side, so the two do not blur.
Genspark's observation volume does not transfer to a one-person company
From the sources
State first what does not transfer. According to the same page, Genspark is “shipping major versions of AI Workspace on a roughly two-month cadence”, and its code comes from “roughly 50 engineers” who “produce all of the company's code through AI tools”.
A&A perspective
The observation volume and the improvement speed are different. A company shipping a major version every two months can observe the change in its difficulty mix release by release. A one-person company substitutes one observation per quarter over a dozen or so re-tests. You give up frequency and, in exchange, standardize the recording side by hand. The precision is lower, but it still separates the two cases that flat satisfaction cannot.
A&A perspective
Draw a line around the number as well. Zhu's “around 80% for a decade” appears on that page without a source, a measurement method or a defined scope. It is stated as a view of the field, and it cannot be treated as a verified statistic. This article does not use it as numeric grounds; it uses it only as the point that flat satisfaction can have more than one cause. It is not a basis for setting your own target at that level.
Where this way of measuring does not hold
A&A perspective
Both pages cited here are self-description by the subject company and its vendor, not independent audits. Chatbase's “Tripled user adoption since integration” and “60-70% of customers in US/Canada” carry no denominator, baseline date or definition on the page. They cannot be used as a forecast for another company, and they are not used as grounds for the test proposed here.
A&A perspective
Demand moving upward with capability is not the only reason satisfaction stays flat. The price may be too high, competitors may have multiplied, response quality may have slipped, or the improvement may simply not be working. Neither source contains evidence that excludes those at the same time. A declined-request ledger is a tool for isolating the capability axis only; it does not rule the other causes out. If the range is widening and satisfaction still does not move, the price and the competitive side need to be looked at separately.
A&A perspective
The ledger has two limits of its own. First, you cannot create the past. A ledger started today produces its first honest increment three months from now. Second, gaps in the record always skew toward the “capability” class. A request declined because it was hard stays in your memory, while a request declined on unit economics was simply never quoted and leaves no trace. Read a skewed ledger as if it were complete and you will overstate the priority of development. Read the first quarter with that skew assumed.
A&A perspective
Finally, what we have not measured. A&A has not measured the effect of this method on retention, on acceptance of a price increase, or on revenue. This article argues about the choice of instrument, not about predicted results, and it offers no figure for how much to raise a price or how far to widen a target. If three months of ledger still leaves the reasons mixed and you cannot separate them, that is a situation where having someone outside sort them once is faster than proceeding with development.
When an improvement does not move satisfaction, the first thing to doubt is the instrument, not the direction of development. Satisfaction only measures the requests you accepted, and that set moves upward together with your capability. What to measure is whether you can now accept a request you were declining six months ago. The material is a ledger of declined requests in five columns. Each quarter, re-test only the rows whose reason was that capability was short: if they pass, move to a price increase or a wider target; if they do not, that is when you question the improvement's direction; and if the declines sit mostly in the other reasons, fix the questionnaire, the contract clause or the price instead. What the Genspark page states is not a satisfaction target but that users adapt and the rate stays flat; the conclusion that satisfaction therefore cannot stand in for capability is ours. On the Chatbase page the change in scope is visible in the list of what the product now accepts — that page does not discuss satisfaction at all. One thing is doable this week: the next time you decline something, write the line.
Sources & editorial note
Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.
- Genspark's Super Agent orchestrates 150+ tools with Claude
Anthropic · no date shown on page
Accessed 2026-10-05 - Chatbase helps companies deliver instant, personalized customer support with Claude
Anthropic · no date shown on page
Accessed 2026-10-05
AI-assisted editorial production
A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.
Editorial check: 2026-10-05