Skip to main content
Navigation

A&A INSIGHTS

Business & AI strategyFor business owners

Routing models per step: decide by failure visibility, not the total bill

Decide AI service model routing from per-step failure visibility, not the total bill: the two conditions a step must satisfy before it may move to a cheaper model.

AI service operationsmodel routingstep designvariable costquality control
日本語で読む
Scattered documents become an organized comparison and a decision
A conceptual illustration of gathering information, organizing it, comparing conditions and making a decision. Illustration generated with AI

THE STARTING POINT

You may move a step to a cheaper model only when a failure in that step can be detected inside the step, and an undetected failure does not propagate downstream. Rather than working down the bill from the largest line, answer those two questions for each step first, and leave any step that fails either test where it is - even when it is the largest line on the invoice.

The answer: only steps whose failure is visible inside the step

A&A perspective

Whether to assign different models to different internal steps is not decided by walking down the bill from the largest line. It is decided by answering two questions about each step first. One: can a failure in this step be detected inside the step itself? Two: if it is not detected, does it flow downstream? Only steps that contain their own failure detection and do not propagate failures may be moved to a cheaper model. A step that fails either test stays where it is, even when it is the single largest line on the invoice.

A&A perspective

The reason for this order is simple. If you start from cost, you will always start with the most expensive step. But the expensive step is usually the one that reads a long context and produces a judgement, and whether that judgement is correct cannot be determined inside that step. You find out that quality dropped only after the output has travelled downstream and reached the customer. Ordering by cost is very nearly the same as ordering by how hard the failure is to see.

From the sources

There is a published record of a company varying the model by step. On 2026-02-17 Gumloop published an account of its own support operation, and the first lesson it lists is to "Know your models, and use the right models for the right tasks". The company describes the purpose as to "optimize for the best intersection of cost and capability" - framed not as cost alone and not as capability alone, but as the intersection of the two.

Gumloop ↗

The total bill does not tell you which step you may touch

A&A perspective

What appears on the invoice is a single number for the month, or at best a breakdown by model. It does not contain how many times each step was called, how many of those calls failed, or where the failure was discovered. In other words the total bill draws no distinction at all between the steps you may change and the steps you must not. Trying a cheaper model in that state is an experiment that can break the whole pipeline without telling you which step broke.

A&A perspective

When one model drives every step, having no such distinction becomes the normal state. You have never looked at quality per step, so quality can only be discussed as one lump called "the quality of the service". While it remains a lump, you also cannot tell which part of it may be lowered. So the first piece of work is not comparing models. It is enumerating the steps and writing down, for each one, how its failures become visible.

A&A perspective

A "step" here means a unit with a clear input and output - one where, if you stopped after it, you could say what gets handed to the next unit. Classifying an incoming request, extracting the required fields, searching the records, inferring the cause, drafting a reply, sending it. Split to roughly that granularity and you can see which steps are self-contained. Conversely, if the only step you can write down is "handling enquiries", you are not yet in a position to decide anything about splitting.

Move or leave, decided from how failure becomes visible in each example step (the judgements are illustrative under stated assumptions, not measurements of ours)
Example stepHow failure becomes visibleDecision
Classify the incoming requestOutput is one of a fixed set of categories; anything outside the set is a fail stated in-stepMove. Record calls, tokens, and out-of-set count
Extract the required fieldsRequired fields and types are fixed; an unfilled field is a fail on the spotMove. Record the count of unfilled fields
Search the recordsA missed record is invisible in-step and flows on as the next step's inputLeave. First build a condition that detects a missed search
Infer the causeAn incorrect inference cannot be judged in-step and reaches the customer downstreamLeave. Do not move it even as the largest cost line
Draft the replyIf a human always reads before sending, that check is the step's inspectionMove if reading happens; leave if it is sent unread
Monitor anomalies and alertWhether the condition is met is decided in-step, and nothing happens when it is notMove. First write both the escalation and the non-escalation condition

Source: Gumloop assigns by the character of the step

From the sources

Gumloop's article is specific about applying different models according to the character of the step. The passage beginning "For lightweight, high-volume tasks, we use Gemini Flash" goes on to assign separate models to code-related work, to core reasoning, and to the hardest tasks. The axis of assignment is not which model is better. It is a property of the step: whether it is lightweight and high-volume, whether it handles code, whether it is core reasoning.

Gumloop ↗

From the sources

The same article also shows that human involvement is positioned differently in different steps. On platform health monitoring it states: "If (and only if) a finding is actionable, it alerts a human". The monitoring itself keeps running automatically, and the condition for reaching a human - that there is an actionable finding - is written into the step rather than left to judgement.

Gumloop ↗

From the sources

It is worth noting that this is the record of a small team. The company writes that "our support team has only two people", explaining that this operation is run by two people. There is no need to read it as something that only works once you have a large dedicated team.

Gumloop ↗

A&A perspective

What transfers from here, though, is the axis of assignment and not its result. The model names and pairings in the article are that company's configuration as of February 2026, not a recommended mix; versions and prices change. And the company does not state the test it applied - what condition a step had to satisfy before being moved to the lighter side. The two conditions below are our attempt to put that test into words.

Two conditions: detectable in-step, and non-propagating

A five-stage flow diagram running left to right, headed "Only a step whose failure is visible here can move to a cheaper model". Stage one is "Cut into steps", with the note "down to units with a fixed input and output". Stage two is the test "Can pass or fail be stated in-step?", with the note "cannot state it, go to the next test". Stage three is the test "Does it propagate downstream?", with the note "it propagates, leave alone". Stage four is the human-check gate "Does a human always read it?", with the note "sent unread, leave alone". Stage five is the conclusion "Move to a cheaper model", with the note "record calls, tokens, in-step fails". The point of the figure is that the order of the tests has nothing to do with the size of the cost, and that a step failing either test or the gate is left alone rather than moved. In other words, even the largest line on the invoice is not moved if its failure is not visible inside the step. That different models are assigned according to the character of each step, with the purpose described as the intersection of cost and capability, and that the monitoring step alerts a human only when there is an actionable finding, both come from the account of its own support operation that Gumloop published on 2026-02-17. That evaluation collects the number of tool calls, token consumption and tool errors as well as top-level accuracy comes from the tool-design article Anthropic published on 2025-09-11. The two test conditions (detectable in-step, non-propagating), the order in which they are applied, and the placement of "leave alone" as a conclusion are A&A's design proposal; neither article states them. This figure shows no cost-reduction rate or quality impact measured by A&A.

A&A perspective

The first condition is that a failure can be detected inside the step. Detectable means you can state pass or fail by looking at the output. Steps whose output has a fixed shape satisfy this. For classification, is the result one of the defined categories? For extraction, are the required fields filled and do the types match? Pass or fail can be stated on the spot, without waiting for the next step.

A&A perspective

The second condition is that an undetected failure does not propagate downstream. Propagating means the error from this step becomes the input of a later step, and the later step then operates correctly on the basis of that error. This is exactly where a summarise-and-pass-on step is dangerous. If one fact drops out of the summary, the summary still reads as a summary. The downstream step treats the incomplete summary as valid input and produces a well-formed error. The more well-formed the error, the later it is found.

A&A perspective

The two conditions are applied in order. If the first is satisfied, the second effectively does not arise, because you can stop the work in that step. For a step that fails the first condition, you look at the second. If it cannot be detected and it flows downstream, leave it alone. If it cannot be detected but a human always reads the output before it flows on, then that human check acts as the step's inspection and the step can be moved. The premise, as in Gumloop's monitoring example, is that the condition for reaching a human is written into the step.

A&A perspective

The effect of these two conditions is to break the cost ordering. It is entirely normal for the most expensive step to end up untouched while cheap, high-frequency steps become the candidates for moving. The saving can look unsatisfying, but you now have an answer about how far you can go without losing quality. Beyond that range the question is no longer which model to buy; it is whether to change the design of the step.

Source: collect calls, tokens and errors as well as accuracy

A&A perspective

Even once you have decided to split on those two conditions, if you have no numbers to look at afterwards, the only way you will learn that quality dropped is from a customer telling you. Having decided per step, you have to hold the records per step too. On what to record, there is a primary account in the adjacent context of tool evaluation.

From the sources

In an engineering article on tool design (published 2025-09-11), Anthropic recommends collecting metrics beyond top-level accuracy during evaluation. The items listed are the runtime of individual calls and tasks, plus "the total number of tool calls, the total token consumption, and tool errors". Accuracy collapses into a single number; these are numbers you can hold separately for each step.

Anthropic ↗

From the sources

The same article states, on judging results, that "Each evaluation prompt should be paired with a verifiable response or outcome". It also recommends that when an error occurs - during input validation, for example - the response should communicate specific, actionable improvements "rather than opaque error codes or tracebacks".

Anthropic ↗

A&A perspective

The scope of this source should be stated plainly. The article is about evaluating tools, not about routing models. Using these items as per-step decision material is our own reinterpretation, not something Anthropic says. It is nevertheless worth reinterpreting, because what the two conditions demand and what this list offers coincide. "Detectable inside the step" means there is a verifiable pass or fail, which is very nearly the same thing as being able to count errors per step.

Write four columns for each step

A&A perspective

Put the decision on one sheet. List the steps down the page and write four columns across. The first is the step name. The second is "can a failure be seen inside this step", and where it can, write concretely what you look at to state pass or fail. The third is "does it flow downstream undetected". The fourth is the conclusion: move or leave, together with the items you will record for that step.

A&A perspective

The important thing about this sheet is that the steps you leave alone stay on it. If the reason for leaving a step alone is not written down, then in a few months the version of you looking at the invoice will want to touch that same step again. If the sheet says "cannot be detected here, therefore left alone", you will know that the next thing to consider is not a cheaper model but re-cutting the step so that detection becomes possible.

A&A perspective

The items to record can be the same three for every step: the number of calls, the tokens consumed, and the number of times the step's own check returned a fail. For a step you moved, you watch the third to confirm it does not move. For a step you left alone, you watch the first two to confirm whether it is even large enough to be worth moving.

For steps you cannot split, define the escalation condition

A&A perspective

A step marked "leave alone" does not have to be left untouched. What you can do in that step is not to change the model but to create a place where failure stops before it flows downstream. Concretely, you write an escalation condition into the step. Writing a condition means deciding two things at once: what has to happen for it to reach a human, and that nothing reaches a human when nothing happens.

From the sources

As a way of phrasing the condition, the wording from Gumloop's monitoring step quoted earlier is instructive: "If (and only if) a finding is actionable, it alerts a human". The condition for escalation is restricted to there being an actionable finding; nothing else reaches a person. Deciding when to escalate is the same piece of work as deciding when not to.

Gumloop ↗

A&A perspective

And once a condition is written and a human always reads the output, that step begins to satisfy the second of the two conditions. In other words it may move from "leave alone" to "candidate". The order cannot be reversed: build the inspection first, change the model second. The substance of the work that lowers your bill is usually not selecting a model but installing an inspection.

Hypothetical: two of six steps turned out to be movable

Hypothetical example

The following is a hypothetical scenario built on stated assumptions. It is not a real price list, not a measurement of ours, and not a customer result. Suppose a solo-run AI service handles enquiries in six steps: (1) classify the incoming request, (2) extract the required fields, (3) search the records, (4) infer the cause, (5) draft the reply, (6) send it. Assume the monthly call counts are ten thousand each for (1) and (2), eight thousand for (3), and two thousand each for (4), (5) and (6). Assume (4) is the step that reads a long context, and that it is also the largest line by token consumption.

Hypothetical example

Apply the two conditions. Step (1) outputs one of a fixed set of categories, so anything outside the set is a fail stated inside the step. Movable. Step (2) has fixed required fields and types, so an unfilled field is a fail on the spot. Movable. In step (3) a missed record is not visible within the step and flows on as the input to (4). Leave alone. In step (4) an incorrect inference cannot be judged within the step and travels through (5) and (6) to the customer. Leave alone. Step (5) is inspected by the human who reads and sends it - but only if the operation actually reads before sending; if it does not, leave it alone. Step (6) does report success or failure in its response, but it does not use a model at all.

Hypothetical example

The result is that two steps, (1) and (2), are movable, and together they account for roughly sixty percent of the monthly calls. Step (4), the largest line, is left alone. The step that cost-ordering would have made you touch first is the one this sheet never moves. And for step (5), a piece of work comes first: confirming whether "a human always reads it" is actually true. If the operation does not read before sending, the conversation is about building that reading step, not about moving the model.

A&A perspective

What this hypothetical deliberately omits is a saving. Prices change by version and the token volume per step changes with the implementation, so putting a figure here would date the assumptions faster than the method. What should be stated is the order of operations: count the movable steps, read their call counts and token volumes out of your own records, and only then look at a price list. Look at the price list first and you are back to talking about the total bill.

Where this does not hold, and the first step today

A&A perspective

The first limit is the maintenance cost of the split itself. Changing models per step adds handoffs and verification, and creates an ongoing need to know which step is running on which model. For a solo operator it is entirely normal for that maintenance cost to exceed the saving. If only two steps are movable and those two are cheap, then not splitting is the correct decision. The two conditions are a tool for granting permission to split, not a tool that recommends splitting.

A&A perspective

The second limit is not yet having cut the work into steps. If the implementation does everything inside one long instruction, per-step decision material does not exist. What is needed then is not a model comparison but cutting the work into steps and deciding each one's input and output. That is a redesign rather than a cost measure, so estimate it as separate work.

From the sources

The third limit is in the sources. On tool design, Anthropic's article states that "More tools don’t always lead to better outcomes" - adding components does not in itself improve the result.

Anthropic ↗

A&A perspective

In addition, Gumloop's article is that company writing about its own operation and Anthropic's is that company writing up its own findings; neither is a third-party audit. No figure for a cost-reduction rate or a quality impact appears in either article, or in this one. Which is to say: nothing here is grounds for splitting steps as such.

A&A perspective

The first step today begins with not opening a price list. Write the steps of your own service on paper, and on each line write one sentence answering "what would I look at inside this step to know it failed?" The lines where you cannot write that sentence are the lines you leave alone. For the lines where you could, pull only their call counts and token volumes out of your records, and see whether that total is large enough to be worth paying the maintenance cost. Compare models after that, not before.

A&A perspective

One clarification: everything above is about assigning models to internal steps. Whether to let the customer choose a model is a separate question with a different decision-maker, covered in "Model selector or fixed default: decide before you ship the setting". Until the internal assignment is settled, you cannot decide what default to present to the customer either.

Whether to assign different models to different steps is settled by how failure becomes visible in each step, not by the invoice. Move only the steps that detect their own failures and do not propagate them, and keep the rest on the same sheet together with the reason they were left alone. For a step left alone, the next thing to do is not to try a cheaper model but to make its failures visible inside the step.

Sources & editorial note

Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.

  1. Supporting the world's most AI-native companies with a 2-person team

    Gumloop · 2026-02-17

    Accessed 2026-10-06
  2. Writing effective tools for agents - with agents

    Anthropic · 2025-09-11

    Accessed 2026-10-06

AI-assisted editorial production

A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.

Editorial check: 2026-10-06

← All articles