A&A INSIGHTS
Do not send AI service support straight to the backlog: fix, explain, build
Fix each destination first, then split support messages into defects, usage gaps and unbuilt use cases before the backlog. From Gumloop's and Chatbase's own accounts.
日本語で読む
THE STARTING POINT
Classify every AI service support message before it reaches the backlog: a defect to fix, a usage gap an explanation can close, or a use case you have not built yet — and assign each class its destination in advance (development, documentation, or a recorded demand note). In practice three more classes join them: a misleading promise, a one-off request, and output drift. Read the next development decision from the distribution across those classes, not from the number of requests.
Split support messages into three classes before the backlog

A&A perspective
When a support message arrives, decide in the same sitting where you write the reply which of three things it is: a defect that reproduces from the same steps; a usage gap, where the product behaves as designed and the customer simply never reached an existing feature; or a use case you have not built, where no path through the current product exists at all. Then fix the destination for each class in advance. Defects go to root-cause identification, usage gaps go to fixing the explanation, and unbuilt use cases go to a demand record. If you stack the customer's own sentences without making that split, the backlog becomes a list of things customers wrote, which is not material for deciding what to build next.
A&A perspective
Classification cannot wait, because the material the decision needs disappears along with the conversation. What had this customer just been doing? Did the behaviour reproduce from the same steps? Was the answer already sitting in the documentation? You know all of that while writing the reply, and you cannot recall it at a backlog grooming session a month later. In a company where one or two people also write the product, intake and development are the same person, so there is no second occasion on which the context gets reassembled. That is why the classification belongs at intake rather than in the monthly pass.
Building the most-requested item turns an explanation problem into code
A&A perspective
Deciding to build in order of request volume has one specific failure, and it is exactly the one that skipping classification creates: messages that only mean the customer could not find an existing feature rise to the top of the count. If the feature exists and ten people still ask the same question, the tally makes it look like the highest-priority development item. But there is nothing to build; what is needed is an in-product hint or a paragraph of documentation. Getting this backwards means spending development time building a second entrance to the same feature, and having two entrances generates its own next wave of questions.
A&A perspective
The mistake also runs the other way. Among the items you classified as solvable by explanation, any where the same question keeps arriving after the explanation is fixed are in fact defects or unbuilt use cases. So the classification does not have to be final on first contact; it is enough to write, per destination, the condition under which the class changes — namely, the same question arriving again after the fix. Counting becomes meaningful only after the classes are attached. Read the distribution instead: say two defects, six explanation gaps and two new use cases out of ten — that tells you directly whether next week goes into development or into writing.
| Class | Distinguishing material | Destination |
|---|---|---|
| Defect to fix | Reproduces from the same steps and the result differs from what the product intends | Identify the root cause as a development candidate |
| Usage gap an explanation can close | The product behaves as designed and the customer never reached an existing feature | Fix the documentation and the in-product hint |
| Use case not built yet | No path through the current product exists at all | Record only the count and the kind of work as demand |
| Misleading promise | Behaviour matches the spec, but the sales material or acceptance conditions invite misreading | Fix the prose that states your scope |
| One-off individual request | Does not apply to any other customer's work | Decide to produce an individual quote, or to decline |
| Output drift, or degradation after a model update | The same steps now give a different output, or the same input gives varying results | Put it into a reproduction set that first settles defect or within-spec |
Gumloop: the kind is decided first, aggregated monthly, then compared with the backlog
From the sources
Gumloop, an automation platform, published an account on 17 February 2026 of how it built its own support operation on its own product, authored by the company's Max Brodeur-Urbas. The piece states plainly that "our support team has only two people". The agent responsible for triaging first determines what kind of problem it is dealing with — "a workflow failure, a how-to question, or an agent issue" — and then, based on that determination, chooses which tools to use next: searching the product documentation for how-to questions, BigQuery for error analysis, or Pylon for edge cases.
From the sources
The connection to the backlog is described as a separate, aggregated, monthly step rather than a per-message one. Every 24 hours an agent fetches the last 24 hours of support tickets and posts a formatted summary to Slack, and then, once a month, "every month a different agent reviews the past 30 daily summaries to extract insights about most common error patterns". The account continues: "This report is compared to our backlog and the agent suggests Linear tickets that should be created to address the most common issues."
From the sources
The how-to class has its own destination, separate from development. A workflow that runs every 24 hours analyses the company's support documentation and, in the account's words, "identifies gaps in our documentation, and uses Devin to draft missing content". Separately, another workflow "categorizes and analyzes every support conversation", which the piece says makes it possible for the support team to identify and understand trends and user sentiment over time.
A&A perspective
What can be extracted from this account is three separations. The kind is determined at the very start of triage. The kind selects the next investigation path, so classification is a branch in the work rather than a tag. And the link to the backlog runs through a monthly aggregate rather than through individual messages. That much is what the source describes; the source places that determination first in triage, and folding it into the moment you write the reply is A&A's own transposition. Moving the same structure to a solo or small company does not require wiring an agent to eighteen tools; it requires a field that records the kind at intake, and an appointment once a month to read the aggregate. That transfer is A&A's reading, not a claim the source makes about small businesses.
Chatbase: spot the trend, then go back to the original conversations
From the sources
On the Chatbase case study Anthropic publishes — a page that lists the company size as Small — founder Yasser Elsaid describes conversation analysis this way: "The AI can summarize exactly what customers want, what parts of the product they dislike, and how their sentiment changes over time." What customers want and what they dislike appear as two separate axes rather than as one satisfaction figure.
From the sources
The same page describes spotting a trend and drilling into the underlying tickets alongside each other: "Managers can quickly spot trends, like multiple customers facing issues with a specific integration, and drill down into relevant tickets to understand root causes." It names the consumer of that analysis as well — "the analytics help product teams identify improvement opportunities". At the same time the page carries Elsaid's remark that "You need humans for the human aspect", and describes the company as streamlining its human-in-the-loop approach so that oversight can happen through simple one-click confirmations.
A&A perspective
Keeping what customers want and what they dislike on separate axes bears directly on how a classification scheme should be built, because collapsing both into a single satisfaction figure adds them together and erases them. A&A's further reading is that the same ordering can be drawn from both sources: look at the aggregated trend first, then return to individual exchanges to check the evidence. The ordering exists to avoid the reverse — generalising from the one message that bothered you. A&A's reading is that the minimum mechanism for holding that order in a small company is one extra line beside the class: the material on which you classified it. The reproduction steps, the screen the customer actually reached, whether a relevant passage already existed in the documentation. That line is what you have to return to when you read the monthly aggregate.
Fix the classes and destinations in a six-row sheet, in advance
A&A perspective
From here on this is A&A's proposal. What the sources name is Gumloop's split into "a workflow failure, a how-to question, or an agent issue" — not the three classes used here. Borrowing that idea and rebuilding it gives three classes, to which practice adds three more, for six rows. The additions are "behaves as specified, but the way the promise was written invites misreading", "a one-off request that does not apply to any other customer's work", and the one specific to a service built on AI: "output drift, or degradation after a model update". The first belongs neither to development nor to documentation; it is a problem with the prose that states the scope of what you sell — the sales material and the acceptance conditions. Mixing the second into the demand record distorts the distribution you are about to read.
A&A perspective
Three columns are enough: the name of the class, the material that distinguishes it, and the destination. The third column is the one that matters, because any class without a destination fixed in advance ends up in the backlog after all. The six destinations are: identify the root cause as a development candidate; fix the documentation and the in-product hint; record only the count and the kind of work as demand; fix the prose that states your scope; decide to produce an individual quote or to decline; and put the case into a reproduction set that first settles whether it is a defect or within spec. Each class maps to exactly one of them. The last one is needed because in a product where the same input can give varying output, the sentence "it does not work" reads equally as a defect and as within spec, and skipping that judgement sends the same case wrongly to both development and documentation.
A&A perspective
This sheet does not automate the decision. The Gumloop account likewise goes only as far as its monthly agent suggesting the tickets that should be created; it does not say who accepts or rejects them, or how. Placing a human decision after that suggestion is A&A's reading, not the source's statement. The sheet's job is that when you read the aggregate once a month, comparable data is still there. Thirty unclassified messages and thirty messages each carrying a class and a line of evidence take the same time to read and yield conclusions of very different reliability.
A hypothetical week: reading the distribution of 24 messages
Hypothetical example
The following is a hypothetical example A&A constructed; it is not a real customer and not a measurement of ours. Assume two people run an AI service with forty contracted accounts. In one week 24 support messages arrive, and sorting them through the six-row sheet yields three defects, eleven usage gaps, five unbuilt use cases, three cases of a misleading promise, one one-off request and one case of output drift. Had you counted only the wording of the requests, the same sentence — "I want to export the list as CSV" — appeared six times, so it would have looked like the highest-priority development item.
Hypothetical example
With classes attached, four of those six turn out to be usage gaps: an export feature already exists and the customer never reached it, so the destination is documentation and an in-product hint. The remaining two are unbuilt use cases, where no path exists in the current product — they want the same export sent automatically each month on fixed conditions — and they enter the demand record. Development time that week therefore goes not to the most-requested CSV item, but to the two of the three defects whose reproduction steps are already confirmed. The single output-drift case goes to neither development nor documentation at this stage; it enters the reproduction set and waits for the defect-or-within-spec judgement. A distribution shaped like this reads as: next week leans towards explanation rather than development. The caution worth stating is that this distribution describes the 24 messages that were written in, not the market.
Where this breaks down, and how to treat the numbers
A&A perspective
There are situations where classification does not pay for itself. While only a handful of messages arrive per week, the founder's memory still holds the context, so the sheet leaves nothing but overhead. It starts to earn its keep once messages outlive that memory, or once a second person begins writing replies. The common objection deserves stating: classifying at intake slows the first response, and sorting once a month in a batch is cheaper. Where intake and development sit with different people that is true, and prioritising response speed is a defensible call. The reason this article lands the other way is that in a one- or two-person company they are the same person, so the "later" in "sort it out later" has nobody else in it. The comparison is between the dozen seconds classification costs and the time spent guessing the context back a month after it was lost. The other limit is that classification and development priority are two different decisions, and this article argues only for the first. A distribution is a frequency among the customers who wrote in; it says nothing about people who left quietly or never signed up in the first place. Not treating the loudest requests as equivalent to market demand is the binding constraint here.
From the sources
The Chatbase page states "Tripled user adoption since integration", but it carries no population, no baseline date and no definition of adoption. The Gumloop account reports its scale as over 500,000 support-related workflows running every week and 18 unique MCP tools, and describes that support operation as built on the company's own product.
A&A perspective
Both are accounts written by the company itself or by its vendor, not independent audits. A&A's judgement is that with no definition written down, the tripling figure cannot be carried over as an expectation for anyone else. The Gumloop scale likewise belongs to a platform company whose support surface is its own product, and a two-person support team inside such a company is not operating under the same conditions as someone carrying the whole business alone. What is taken from it here is the structure — the kind decided first, before the case moves on — and not the numbers.
A&A perspective
It is worth stating plainly what A&A has not measured. We have no measurement showing that this split reduced inquiries, increased closed deals or improved retention. This article is two published firsthand accounts plus a design we reassembled from them for small businesses. Where this step sits in the wider flow is covered in "AI-native GTM: a practical guide for solo founders and small teams". For reducing the load on the reply side, "Repeated questions consume the team: an Intercom-inspired way to create support capacity" handles the adjacent decision, and for setting the conditions to hand a conversation back to a human before AI answers it, see "Before letting AI receive support inquiries, define the conditions for handing them back: measure wrong resolutions, not just answer rate". If you cannot settle internally which class deserves development next, an initial consultation can start from how to read your own distribution.
To stop support messages from piling up as raw feature requests, split them at intake into three classes — defect, usage gap, unbuilt use case — and write the destinations in advance into a six-row sheet that adds the three classes practice produces: a misleading promise, a one-off request, and output drift. Read the distribution once a month and decide whether the coming week goes to development or to explanation. That distribution is a frequency among the customers who wrote to you, not the demand of a market.
Sources & editorial note
Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.
- Supporting the world's most AI-native companies with a 2-person team
Gumloop · 2026-02-17
Accessed 2026-09-28 - Chatbase helps companies deliver instant, personalized customer support with Claude
Anthropic · not stated on page
Accessed 2026-09-28
AI-assisted editorial production
A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.
Editorial check: 2026-09-28