Skip to main content
Navigation

A&A INSIGHTS

Business & AI strategyFor business owners

Agent or Workflow? Count the Steps Before You Quote an AI Build

Agent or fixed workflow? In AI contract work it turns on one question: can you count the required steps before you quote? Countable steps go fixed-price.

AI contract workagentsworkflowsscopingacceptance criteria
日本語で読む
Scattered documents become an organized comparison and a decision
A conceptual illustration of gathering information, organizing it, comparing conditions and making a decision. Illustration generated with AI

THE STARTING POINT

Whether to take a job as an agent or as a workflow with a fixed path turns not on how sophisticated the build is, but on one question: can you count the required number of steps before you quote? Steps you can count go on a fixed path at a fixed price; hand only the uncountable steps to the model, and on those promise the range of permitted tools and the handoff recipient rather than the path itself.

The answer: the dividing line is whether you can count the required steps before you quote

A&A perspective

A customer asks you to build them an agent. You review the requirements and the steps turn out to be largely fixed. There is exactly one thing to decide here: do you take the job as an implementation whose path you write out in advance, or as one where the model decides the path? The test is not how sophisticated the build is. It is the single question of whether you can count the required number of steps before the quote goes out.

A&A perspective

If you can count them, take it as a fixed path. What you are selling in that case is predictability. You can write the input-to-output correspondence step by step, so both the acceptance criteria and a fixed-price quote hold up. If you cannot count them, hand only that step to the model. But the path will differ every time, so the path is not what you can promise. What you can promise is the range of tools it may use, and the rule for who receives the work when it cannot resolve the case.

A&A perspective

And the unit of this judgement is the step, not the project. It is normal for one system to contain both steps whose moves you can read in advance and steps whose moves you cannot. If you classify the whole quote as an agent job or a workflow job, you end up paying the cost of autonomy on the steps you could have read in advance.

Anthropic's distinction: predefined code paths versus the model directing itself

From the sources

Anthropic's "Building effective agents" (published 19 December 2024) separates the two structurally. On workflows: "Workflows are systems where LLMs and tools are orchestrated through predefined code paths." On agents: "Agents, on the other hand, are systems where LLMs dynamically direct their own processes and tool usage, maintaining control over how they accomplish tasks."

Anthropic ↗

From the sources

The same page states the applicability condition: "Agents can be used for open-ended problems where it's difficult or impossible to predict the required number of steps, and where you can't hardcode a fixed path." What the condition turns on is not the difficulty of the task, but the predictability of the number of steps.

Anthropic ↗

From the sources

It is also explicit about how to choose: "workflows offer predictability and consistency for well-defined tasks, whereas agents are the better option when flexibility and model-driven decision-making are needed at scale". On sequencing, it says to "add multi-step agentic systems only when simpler solutions fall short".

Anthropic ↗

A&A perspective

For a contractor, the word that matters in that quote is predictability — because predictability is a property you can write straight into a quote. That connection, however, is not the source's claim. Anthropic is giving design guidance for building systems; it says nothing about contract pricing or acceptance criteria. That predictability is the easiest thing to sell is A&A's reading.

Anthropic ↗

The character of each step changes the quote line and the acceptance check (an A&A design proposal. What the sources establish is this much: that workflows are systems where LLMs and tools are orchestrated through predefined code paths; that agents are systems where LLMs dynamically direct their own processes and tool usage; that agents can be used for problems where the required number of steps cannot be predicted and a fixed path cannot be hardcoded; that agentic systems trade latency and cost for task performance and carry higher costs plus the potential for compounding errors; that extensive testing in sandboxed environments with appropriate guardrails is recommended; and that Gumloop's two-person support team splits workflows and agents per step and alerts a human if and only if a finding is actionable. Treating steps as quote line items, and the content of each row, are A&A's own construction)
Character of the stepLine to write in the quoteWhat acceptance checks
The moves can be counted (writable out, branches included)Input-to-output correspondence, fixed priceWhether prepared inputs yield those outputs
The tools it uses are not settled before it runsThe list of permitted tools, and what happens outside that rangeWhether out-of-range inputs were forced into a classification
How many moves it takes to finish is not settledThe cap on attempts, and the recipient once the cap is hitWhether capped-out cases reached the agreed recipient with the information attached
Failure is not detectable automaticallyThe condition for escalating to a person, and the destinationWhether only cases meeting the condition were escalated
Work common to every autonomous stepSandboxed testing, guardrail design, error detection in operationThe testing record, and whether stop conditions actually bite

On a fixed-price job, the cost of autonomy comes out of your own margin

From the sources

The source is explicit that autonomy has a price: "Agentic systems often trade latency and cost for better task performance, and you should consider when this tradeoff makes sense."

Anthropic ↗

From the sources

It also itemises that price: "The autonomous nature of agents means higher costs, and the potential for compounding errors." Two things — cost, and errors that accumulate.

Anthropic ↗

A&A perspective

In contract work, who pays for those two is already settled before the build starts. If the number of steps grows on a fixed-price engagement, the extra inference cost comes out of your share. If errors accumulate, the work of finding them and putting the job back falls inside operations. The cost of autonomy is therefore not a question of design taste. It is a question of which wallet it comes out of.

Anthropic ↗

A&A perspective

So selling autonomy on a step whose moves you could have counted is not merely over-engineering. It is a trade in which you give up predictability — the easiest value you have to sell — and accept variable cost and operational liability in exchange. The proposal looks more sophisticated. The margin gets worse.

Gumloop's public record: the split was per step, not per project

From the sources

There is a public record of a team drawing the line per step. Gumloop's "Supporting the world's most AI-native companies with a 2-person team" (Max Brodeur-Urbas, 17 February 2026) describes how the company built its own support operation on its own product, and states that "our support team has only two people". The parts used were "the same building blocks that are available to every Gumloop user" — described as workflows, agents, triggers and MCP tools.

Gumloop ↗

From the sources

What sits on a fixed path are the steps whose inputs and order are settled. Per the article, user-data enrichment is a workflow that queries internal databases every time a ticket opens and pushes a complete user dossier in, with that data synced every 30 minutes. Identifying the day's follow-ups, analysing the support docs for gaps every 24 hours, and categorising and analysing conversations are all likewise described as workflows.

Gumloop ↗

From the sources

What the model is allowed to decide are the steps whose moves cannot be read in advance. The diagnosing agent, named "Gummie Support", "first determines what kind of issue it's dealing with: a workflow failure, a how-to question, or an agent issue". From there, "Based on the situation, it chooses which tools to use next".

Gumloop ↗

From the sources

Model selection is per step too: "at Gumloop, we use many different models across our agents and workflows", with different models assigned to lightweight high-volume tasks, code-related tasks, core reasoning and the hardest tasks. The figures carried on the page — over 500,000 support-related workflow runs every week, 18 unique MCP tools, and a support-ticket response time of under five minutes — are all published by the company about itself.

Gumloop ↗

A&A perspective

What a contractor can take from this record is not the quantity of autonomy but the location of the line. Collecting user information once a ticket opens has a settled number of moves, so it sits on a fixed path. Working out what kind of problem this even is does not, so it is handed to the model. The boundary is drawn inside one system; neither option was chosen for the whole thing.

Gumloop ↗

Three questions to ask of each step before you quote

A five-stage flow running left to right. The heading reads "Count the moves, then split the line". Stage one, "Split into steps": break the paragraph of requirements you took down into one line per step. The unit of judgement is the step, not the project. Stage two, "Can you count the moves?": ask whether the required moves can be written out with branches included, and treat the steps that can be as a fixed path. Stage three, "Are the tools settled?": ask whether the tools the step uses are settled before it runs; if they depend on the situation, the choice of tool itself is being delegated to the model. Stage four, "Will you know it failed?": ask whether a failure in this step is detectable; a step where it is not cannot get an acceptance condition until its handoff rule to a person is settled. Between stages four and five a vertical approval gate marks the condition for escalating to a person, annotated with the recipient and the information attached. Stage five, "Quote per step": fixed-path steps go in at a fixed price on an input-to-output correspondence, while steps handed to the model are stood up separately with a cap on attempts. The point of the diagram is that agent-versus-workflow is not chosen per project: the three questions are applied per step to split the lines of the quote. That workflows are systems orchestrated through predefined code paths, that agents are systems where LLMs dynamically direct their own processes and tool usage, that agents can be used for problems where the required number of steps cannot be predicted and a fixed path cannot be hardcoded, and that agentic systems trade latency and cost for task performance while carrying higher costs and the potential for compounding errors all come from Anthropic's "Building effective agents"; splitting workflows and agents per step and alerting a human if and only if a finding is actionable comes from Gumloop's public record. The ordering of the three questions, the placement of the gate and the assignment into quote lines are A&A's design proposal, and the diagram presents no margin or win-rate improvement measured by A&A.

A&A perspective

Split the paragraph of requirements you took down into steps, then ask the following three of each step in turn. These are the source's applicability condition restated as quoting work.

A&A perspective

First: before this step runs, can you count the moves it needs? Not "roughly three" — can you write them out, branches included? If you can, it is a fixed path. If you start writing them out and realise that it is not settled until you have seen the input, that is where the boundary is.

A&A perspective

Second: are the tools this step uses settled before it runs? If they are, you can also write the order. If which tool is needed depends on the situation, then you are delegating the choice of tool itself to the model. That is precisely where Gumloop's diagnosing step above is split off.

Gumloop ↗

A&A perspective

Third: when this step fails, will you know it failed? On a fixed path you know which move it stopped on. On a step where the model owns the path, it can run all the way to the end with a wrong judgement in the middle. Any step that produces an "I would not know" here cannot get an acceptance condition until you have settled its handoff rule to a person first.

The character of the step changes the line you write in the quote

A&A perspective

The answers to those three questions change what goes into the quote. For a fixed-path step, write the input-to-output correspondence step by step. Acceptance can be checked by asking whether this input yields this output, which is what makes a fixed price defensible.

A&A perspective

For a step handed to the model, you cannot write the path. Write three other things instead. The list of tools it may use. The range in which it may resolve the case itself. And who receives the case, with what information attached, when it cannot. With those three settled, something remains to accept even though the path differs every time.

A&A perspective

It also matters not to collapse a step that straddles the boundary into one line. Write "receives enquiries and responds appropriately" and you have put the information-gathering part (settled number of moves) and the what-kind-of-problem-is-this part (not settled) on the same line, after which neither the price nor the acceptance check can be fixed. Splitting the lines is the quoting work.

From the sources

On not taking the whole thing at once, the source side supports the point. Gumloop's article says the support system "didn't start out fully formed: it grew one workflow at a time, over months of continuous iteration". Even a system at the scale two people can run was not designed in a single pass.

Gumloop ↗

On a delegated step, what you promise is the handoff condition, not the path

From the sources

The same article contains a worked instance of what can be promised on a step handed to the model. Of the agent watching platform health, it says: "an agent automatically investigates error rates and monitors platform health. If (and only if) a finding is actionable, it alerts a human."

Gumloop ↗

From the sources

The social-monitoring step takes the same shape. Per the article, an agent monitors Reddit and X/Twitter for relevant mentions of Gumloop, detects which mentions are customer complaints, and responds to those complaints. The article states where it proceeds to: replying to the mentions it judges to be complaints.

Gumloop ↗

A&A perspective

What is fixed in both of these is not the route of the investigation or the monitoring. It is a condition and a destination: escalate to a person only when the finding is actionable; reply to the mentions judged to be complaints. How it investigates its way there is not described in the article. The route may differ every time. The promise holds because the condition and the destination are fixed.

Gumloop ↗

A&A perspective

Contract acceptance criteria can be written in the same shape. Not "the investigation follows the same procedure every time", but "when this condition is met, the case reaches this recipient with this information attached". How to turn output variance itself into acceptance criteria is a separate decision, and that is covered in "Acceptance criteria for AI deliverables that vary from run to run".

Three cost lines that only the autonomous steps carry

From the sources

The autonomous side carries work the fixed-path side does not. As a countermeasure, the source states: "We recommend extensive testing in sandboxed environments, along with the appropriate guardrails." Preparing a sandboxed environment, testing extensively and designing guardrails are all work somebody performs.

Anthropic ↗

From the sources

The nature of the cost also differs. The same page says that agentic systems trade latency and cost for better task performance, and that the autonomous nature of agents means higher costs. More moves means more inference calls. On a fixed path the number of moves per case is settled, so the cost per case can be estimated too.

Anthropic ↗

A&A perspective

So in the quote, stand up three separate cost lines against the autonomous steps. Testing in a sandboxed environment. Designing the guardrails — the restriction on usable tools, the stop conditions, the cap on attempts. And, after launch, the work of finding accumulated errors. None of this needs to be loaded onto the fixed-path steps; load it on and you simply have an expensive quote.

Anthropic ↗

A&A perspective

Cost per case can be folded into a fixed price for a fixed-path step. On the autonomous side the number of moves per case varies, so if you are going to fold it in, decide up front either a cap on attempts or a separate usage-based line. Either is fine. Put it into a fixed price without deciding, and you absorb the overrun in the months it gets used heavily.

A hypothetical example: splitting "build me an agent" into six steps

Hypothetical example

Here is a hypothetical example, split into steps. The situation: a customer says they want an agent built to take internal enquiries. On reviewing the requirements, enquiries arrive in a single chat channel, their content falls almost entirely into three kinds — checking a policy, an expenses procedure, and a device fault — and the policies and procedures already exist as documents.

Hypothetical example

Split into steps, there are six. Receiving the enquiry (moves settled). Retrieving the poster's department and permissions (settled). Judging which of the three kinds it is (not settled until the input is seen). Presenting the relevant passage for a policy or procedure (settled). Isolating the cause for a device fault (not settled). Handing off when it is not resolved (settled). Four fixed-path steps, two handed to the model.

Hypothetical example

The quote gets one line per step. The four fixed-path steps get their inputs and outputs written out, at a fixed price. The judging step gets the information it may use (the body of the post and the poster's department) and what happens when the enquiry is none of the three (route it to handoff rather than classify it). The isolating step gets the tools it may use (the device inventory and past response records), how many attempts it may make, and who receives the case with what attached when it is not resolved.

Hypothetical example

Acceptance divides by line too. The four fixed-path steps are checked by confirming that prepared inputs yield those outputs. The judging step is checked not on its route but on whether it refrained from forcing enquiries outside the three kinds into one of them. The isolating step is checked on whether cases that exhausted the cap without resolving reached the agreed recipient with the agreed information.

Hypothetical example

It is worth writing down what happens without this split. Take all six steps as a single line — "an agent that resolves enquiries autonomously" — and the fixed price now contains two steps whose number of moves is not settled until the input arrives. In heavily used months the inference cost rises, and cases where the isolation was wrong get picked up in operations. Both happen outside the quote.

A&A perspective

This example was constructed for explanation. It is not an A&A engagement, nor a record of anything delivered in this shape. In a real engagement the number of steps and the position of the boundary both move with the share of enquiries that fall outside the three kinds, and with whether the device inventory may be touched at all.

Where this test does not apply, and what is not being claimed

A&A perspective

Four conditions where this test does not apply.

A&A perspective

First, a fixed path can fail to pay even when the moves are readable. What the source separates is the predictability of the number of steps, not the rate at which the inputs change. If the customer's upstream — the channel enquiries arrive through, the format of a form — changes every month, a fixed path becomes a monthly change request. That is not a reason to sell autonomy, though; it is a reason to price the change path up front. That connection is A&A's reading.

Anthropic ↗

A&A perspective

Second, real engagements are not a clean binary. The common shape is a fixed path with a single model-directed step inside it, and the only way to quote that is line by line per step. The three questions here are a tool for splitting lines, not for classifying projects.

From the sources

Third, Gumloop's record is the company's account of using its own product on itself, not an independent third-party audit. The published 500,000-plus weekly runs, 18 unique MCP tools and sub-five-minute response time are figures the company published about its own scale and operation. And the fact that two people could run it depends on that support work being contiguous with the product they build: as the article itself says, "Everything described above was built using the same Gumloop platform available to every customer" — the system in question is the product. None of it is a forecast for a small contracting business, or an attainable benchmark.

Gumloop ↗

From the sources

Fourth, Anthropic's text is design guidance from a model provider. Published 19 December 2024, it gives a sequence — "add multi-step agentic systems only when simpler solutions fall short" — but it does not address contract margin, pricing method or acceptance criteria. Every restatement into quotes and acceptance in this article is A&A's reading.

Anthropic ↗

A&A perspective

What is not claimed: this article carries no A&A engagement record, win rate or margin, and no price level. The hypothetical example is a construction for explanation, not a delivery record. No causal claim is made that splitting steps improves margin. The claim extends only this far — that splitting them lets you see, before the quote goes out, which steps are selling predictability.

The next step: split one live requirement into steps and mark the moves on each line

A&A perspective

The next step can be small. Take the one engagement you are about to quote, split the requirements paragraph into steps, and write nothing on each line but whether you can count its moves. If you can write yes on every line, that engagement does not need to be taken as an agent at all.

A&A perspective

If a line comes out as no, write two more things on that line only: the tools it may use, and the recipient when it cannot resolve the case. A line where you cannot write those two is not yet in a quotable state. Either there is still something to ask the customer, or the business process itself has not been decided.

A&A perspective

If you cannot settle how to split the lines at all, confirming the work comes first. What to settle before starting so that rework shrinks is covered in "What to settle before an AI implementation quote", and who takes on what after delivery is covered in "After an AI workflow launches, who restores the interrupted work?". If the split is settled and you want to fix the scope of the build, an A&A development conversation is the right fit. For the whole sequence first, there is "AI-native GTM: a practical guide for solo founders and small teams".

Whether to take a job as an agent is not a judgement about sophistication. It is the work of splitting the requirements into steps and checking, before the quote goes out, whether the moves in each step can be counted. Sell autonomy on a step you could have counted and you hand back predictability — the easiest value you have to sell — and take on variable cost and operational liability in its place.

Sources & editorial note

Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.

  1. Building effective agents

    Anthropic · 2024-12-19

    Accessed 2026-10-03
  2. Supporting the world’s most AI-native companies with a 2-person team

    Gumloop · 2026-02-17

    Accessed 2026-10-03

AI-assisted editorial production

A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.

Editorial check: 2026-10-03

← All articles