A&A INSIGHTS
Vertical AI vs. generic chat: sell the scope of context you maintain
Vertical AI service founders: decide which of approved wording, brand assets, usage occasions and check criteria you take on, and which stay with the customer.
日本語で読む
THE STARTING POINT
What separates a vertical AI service from a general chat assistant is not access to a model but which parts of the work context it keeps usable over time, and who answers when one of them is wrong. Put on one sheet which of the four objects — approved wording, brand assets, usage occasions, check criteria — you take on and which stay with the customer; only then do you have a reason to be paid that is not a model name.
The answer to "couldn't we just use a general chat assistant?" is the scope of context you take on
A&A perspective
When a customer asks whether a general-purpose AI chat assistant would be enough, the answer is not a difference in model quality. It is the scope: which parts of the work context your service keeps usable over time, and who answers when one of them turns out to be wrong. For a vertical AI service in brand operations, that scope splits into four managed objects — approved wording, brand assets, usage occasions, and the check criteria. Once you can put on a single sheet which of the four you take on and which stay with the customer, you finally have a reason to be paid that is not a model name. Until you can write that sheet, the customer's objection is the more accurate one.
A&A perspective
The reason is simple. Access to a model is something the customer can buy directly. A large context window, and the act of feeding a long document into it, are both available on their side of the desk. What customers do not want to do themselves is keep track of which version of an asset is currently in force, decide which materials get consulted for which kind of work, judge whether an output met the standard using the same method every time, and explain the call when a judgement turns out to be wrong. None of that is a capability question; it is a question of responsibility carried continuously. So the product does not take the shape of "a more capable AI". It takes the shape of a promise: these parts are ours.
A&A perspective
This article uses two overseas primary sources. One is the brand.ai customer story, which describes a vertical service for brand operations and states that the company built two separate mechanisms for handling context. The other is Anthropic's engineering write-up on retrieval, which describes what is lost when material is cut into fragments, and at what size the retrieval machinery is not needed at all. Both are written by the publishing organisation itself and are not independent audits. Throughout, source statements, A&A's own reading, and the fictional worked example are kept in separate blocks.
brand.ai: finding the brand insight and holding the conversation are built separately
From the sources
The brand.ai customer story published by Anthropic states that the company built two complementary retrieval systems around Claude: "They implemented two complementary retrieval systems around Claude—one for surfacing relevant brand insights, another for maintaining context in conversations." One exists to surface what is known about the brand at the moment it is needed; the other exists to hold context across the exchange.
From the sources
The same page names the size of the context window as one reason for its model choice, describing "its industry-leading context window that allows it to process entire brand guidelines at once." Handling the guideline as a whole object, rather than in pieces, is stated as a condition of the choice.
From the sources
The page also names what is being displaced: "The traditional approach of static brand guidelines created by agencies no longer works in a world of constant content creation across global teams." Co-founder Chelsey Susin Kantor describes the resulting state of play as one where "Teams end up with a culture of 'it's good enough' because they simply don't have the time to do better," with the team's attention going to policing brand compliance rather than growing the brand.
A&A perspective
The useful part here is not any outcome figure. It is the structural fact that the mechanism is split in two. Surfacing the right knowledge and holding the premise mid-conversation are different jobs even when they read the same material. The first can be working perfectly while the tone specified at the start quietly drops out by the third exchange — and from the user's seat, that output is a failure. Holding the assets and still honouring the same premise several turns later do not come together unless they are designed separately. Translated into the language of an offer: "we hold your materials" and "we do not lose them mid-conversation" are two different promises.
A&A perspective
The same page also carries figures — an annual brand-compliance cost reduction and a per-copywriter content volume among them. Those are vendor statements published by Anthropic, with no method given and no independent verification, and the stated subject is "for enterprise clients". They cannot be used as a forecast for a small customer, so this article does not rest any claim on them. What it uses is the structural fact that two mechanisms were built separately.
| Managed object | What the service promises | What stays with the customer |
|---|---|---|
| Approved wording (permitted phrasings and prohibited expressions) | Hold the list and flag any output falling outside it for return | Deciding what goes on the list, and approving additions and removals |
| Brand assets (guidelines and past work) | Show which version is in force and reference it; on replacement, produce the list of work made against the old one | Licence to use the material, and notifying you when a version changes |
| Usage occasions (who produces what, in which channel) | Vary the materials consulted and the strictness of the check by occasion | Declaring when a new channel or department is added |
| Check criteria (pass or fail) | Publish the items judged on and attach the reason for a failure | The final call when they disagree with a judgement |
| Conversational context (continuity with the same person) | Hold the opening premise through the exchange and avoid carrying stale premises forward | How long and how much may be retained |
| Generation itself (writing the copy) | Not promised here. It is the model provider's capability and the customer can buy it directly | The final decision on whether to use an output |
Anthropic on retrieval: a fragment stops being usable the moment its surroundings are cut away
From the sources
Anthropic's Contextual Retrieval post, published on 19 September 2024, states the problem with ordinary RAG directly: "traditional RAG solutions remove context when encoding information, which often results in the system failing to retrieve the relevant information from the knowledge base." The context is lost at the moment the material is cut into chunks and encoded, and the needed information then cannot be pulled back out.
From the sources
The post gives a worked example. The fragment "The company's revenue grew by 3% over the previous quarter." cannot say on its own which company or which period it refers to — "this chunk on its own doesn't specify which company it's referring to or the relevant time period." The remedy described is to prepend explanatory context to each chunk before creating the embedding and the index, with that added context usually running 50 to 100 tokens.
From the sources
The effect is measured on Anthropic's own evaluations. Combining contextual embeddings with contextual BM25 reduced the top-20-chunk retrieval failure rate by 49% (5.7% to 2.9%), and adding a reranking step reduced it by 67% (5.7% to 1.9%). The post states the bounds of that measurement too. The domains were "codebases, fiction, ArXiv papers, Science Papers", the evaluation used the question-and-answer sets they themselves used for each domain, and the metric is given as "We use 1 minus recall@20 as our evaluation metric" — the share of relevant material that fails to appear in the top 20. The reported figures are for "the top-performing embedding configuration (Gemini Text 004)".
A&A perspective
Move this to brand operations and the same problem appears outside the retrieval layer too. A line that says "avoid this phrasing" cannot be applied to an actual piece of work unless it also carries which product, which channel, and from what date the decision holds. That is exactly where the assumption "we'll just paste the guideline PDF in" breaks down. Pasting puts the whole document in front of the model, but the work of choosing which lines are in force for this particular job stays with a person. One caution, stated plainly: the 49% and 67% figures describe retrieval accuracy, not brand-copy quality, and they were measured on a different kind of material. There is no basis for promising a customer the same improvement.
Split the managed context into four objects, then decide which ones you take

A&A perspective
What the sources establish stops at two structural points: surfacing and holding are different jobs, and a fragment needs its surroundings described. Everything from here is A&A's construction. Divide the context a vertical brand-operations AI could manage into four objects, treated as units of responsibility: approved wording (permitted phrasings, prohibited expressions, the correct form of a term), brand assets (guidelines, past work, which version is currently in force), usage occasions (who consults which materials when producing what, in which channel), and check criteria (the items an output is judged on, and how a failure is explained).
A&A perspective
Laid out that way, the sales conversation becomes legible. "Wouldn't a general chat assistant do?" is really the question "which of these four do you hold?" If the answer is none of them, then there genuinely is no difference from the customer pasting their own guideline into a chat window. If the answer is all four, you have taken on final authority over the brand's wording and you will collide with the customer's own brand owner. Differentiation is not taking everything. It is drawing the boundary before the customer draws it for you.
A&A perspective
In practice the two easiest to take on first are brand assets and usage occasions. Both involve concrete work whose result the customer can verify. For assets: "we hold which version is in force, and when it is replaced we produce the list of published work made against the old one." For occasions: "advertising, recruiting and in-store each consult different materials at different levels of strictness, and you tell us when a new channel appears." Approved wording touches the customer's own authority to approve, so the workable split is to hold the list and flag anything outside it for return, while the decision about what goes on the list stays with the customer. Check criteria are the heaviest of the four: the moment you take them, you have also taken responsibility for explaining a wrong pass.
A&A perspective
Write it as three columns: the managed object, what the service promises, and what stays with the customer. The table at the foot of this article has that shape. When you put it in a proposal, add a fourth column for how the customer verifies each row. Go row by row and ask what evidence would show the promise held; any row you cannot answer for should come off the promise. What stays is only what the other side can check for themselves. Organising your own internal business context is a separate question, covered in "Prepare reusable business context before repeating the same AI briefing" — that one is about tidying context for your own use, while this article is about deciding what scope you take on for a customer. If what you want is a comparison of whole service shapes instead, that is "AI-native service-company cases: what Advolve, Newfront and Lindy help founders compare".
"We built a retrieval system" is not, by itself, a product
From the sources
The same Anthropic post is explicit about when the machinery is unnecessary. If a knowledge base is "smaller than 200,000 tokens (about 500 pages of material)", the whole thing can go into the prompt "with no need for RAG or similar methods". The post states that threshold in its own terms, as roughly 500 pages of material.
A&A perspective
That description dates from September 2024, so it does not account for later model versions or pricing; what is borrowed here is the shape of the finding — below some volume the machinery is not needed — rather than the number itself. A small or mid-sized customer's brand guideline will usually not reach 500 pages of material. Which means that if "we run a vector search" is your differentiation, you lose the argument to a customer who pastes the whole document in themselves. The presence or absence of retrieval machinery is a condition that starts to matter once the volume crosses a threshold; it is not a reason to be paid. If the differentiation box in your proposal currently contains only a technology name, that is the box to rewrite.
From the sources
Among its implementation considerations the same post lists "Always run evals", and in that item states: "Response generation may be improved by passing it the contextualized chunk and distinguishing between what is context and what is the chunk." Passing the added context and the chunk itself as distinguishable things may, it says, improve what gets generated.
A&A perspective
Read as an instruction about an offer, "always run evals" means the check criteria are a designed object rather than an inspection bolted on afterwards. Of the four managed objects, check criteria are the heaviest precisely because they are the only one that generates continuous work. Assets might be swapped once a month; a judgement runs on every output. The other side of that is the reason a customer keeps renewing: going back to a general chat assistant means taking back the job of deciding pass or fail every single time. That is a reason to stay, not an inability to leave.
A&A perspective
So what belongs in the sales material is not the name of the technology but the answers to: which materials, updated by whom, on what cycle, and against what criteria an output passes or fails. None of that is replaceable by a general chat assistant. What is replaceable is the act of pasting a document in and generating once. If that single act is what you are selling, the price only goes down.
A fictional example: answering "why not just use ChatGPT?" on one page
Hypothetical example
What follows is a fictional situation constructed by A&A. It is not a real customer and not a delivery record. Suppose the communications lead at a 40-person cosmetics manufacturer, hearing about a monthly brand-copy checking service, says: "couldn't we just paste the guideline into our in-house AI chat?" Answering with a feature list fails here, because each feature individually collapses into "a chat assistant can do that too." Instead, hand over one page with the four managed objects as its rows.
Hypothetical example
The brand-assets row reads: "we hold which version is currently in force, and when it is replaced we produce the list of published work that was made against the old one." The usage-occasions row: "advertising, recruiting and in-store consult different materials at different levels of strictness, and you tell us when a new channel is added." The approved-wording row: "we hold the list of prohibited expressions and return anything outside it; what goes on the list is your decision." The check-criteria row: "we publish the items an output is judged on and attach the reason for any failure; where you disagree, your communications lead decides." A fourth column on every row states how the customer can verify it.
Hypothetical example
If they say they will manage versions themselves, that row goes back to them and the monthly fee comes down. If they want the judgement handed over as well, you first tell them how many items per week you can actually review, and then take it. The purpose of the page is not persuasion; it is to let the other side choose how much of it they are buying. If all four rows go home with the customer, then a general chat assistant really is enough for them, and declining the contract costs less than the cancellation later.
A&A perspective
Note that the example contains no amount of money and no percentage saved, because nothing about the starting state has been measured. If you are going to put a number in front of a prospect, decide first what will be counted. For instance: the number of pieces sent back before publication, the number of published pieces found to have been made against a superseded asset, and the number of objections raised against a judgement. Count those by hand for one month and set them beside the same count after the start. All three are things the customer can count themselves. A saving you have not measured will have to be explained at the first renewal meeting.
Where this explanation does not hold, and how to check
A&A perspective
There are three cases where it does not hold. The first is a customer whose material is small to begin with and rarely updated. The threshold discussed in the previous section applies directly: passing the whole thing in is enough. For that prospect there is no row among the four you can meaningfully take. Deciding they are out of scope is faster than chasing the contract.
A&A perspective
The second is when you do not have the time to run the checking work yourself. Taking on the check criteria brings with it the responsibility for explaining a wrong pass. If you run the business alone and cannot say how many items per week you can review, do not take that row. Estimating how much human review a service can absorb is covered separately in "Before taking on more AI delivery work, estimate where human review jams."
A&A perspective
The third is when you do not have permission to use the material. Ingesting a customer's past work, or third-party rights-holding material, for reference requires agreement on scope. Skip that and the better your system gets, the harder it stops later. Related to this: the hope that accumulated context will make a customer unable to switch is not guaranteed here either. What accumulates is the customer's asset, not yours. The thing that actually sustains the relationship is that you keep carrying the updating and the checking — work that shows up every month.
A&A perspective
Here is how to check. Take your three most recent sales conversations and write down, in one line each, the last thing the other side asked about or wanted confirmed. Classify each line against the four managed objects. If none of them fits — if all you can write down is model names and generation speed — then this design has not yet entered your proposals. If all three cluster on the same row, that row is your product as of today, and that is the row whose evidence you should design first. For the wider picture from acquisition through retention, see "AI-native GTM: a practical guide for solo founders and small teams."
What answers "wouldn't a general chat assistant do?" is not a description of a model but the scope of context you take on. Put the four objects — approved wording, brand assets, usage occasions, check criteria — on rows, and write out what the service promises, what stays with the customer, and how the customer can verify each one. Any row you cannot write comes off the promise. Start this week by taking your three most recent sales conversations, writing down in one line each what the other side last asked to have confirmed, and classifying those lines against the four. If most of them will not classify, what you are selling is still access to a model.
Sources & editorial note
Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.
- Brand.ai uses AI to make brands more human with Claude
Anthropic · no publication date shown on page
Accessed 2026-09-28 - Introducing Contextual Retrieval
Anthropic · 2024-09-19
Accessed 2026-09-28
AI-assisted editorial production
A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.
Editorial check: 2026-09-28