A&A INSIGHTS
Should AI research split across multiple agents? Test the question first
What decides whether to parallelise research is the shape of the question, not model capability: does it split into directions that need no results from each other?
日本語で読む
THE STARTING POINT
Whether to split research across multiple agents is decided by the shape of the question, not model capability: only directions that need no intermediate results from each other gain from splitting. Price the integration and per-direction review before deciding.
The answer: the question's shape decides, not the model
A&A perspective
Whether to split research across parallel agents is decided by whether the question you accepted decomposes into directions that do not need each other's intermediate results. If it does, division of labour pays. If it does not, adding agents adds duplication and gaps. The order of the decision is three steps. First, break the question into directions and write, in one line per direction, whether that direction is independent. Second, write delegation conditions for the independent directions only. Third, put the review time for the number of directions you split into, plus the integration effort, into the estimate. There is no reason to change the architecture before those three are written.
From the sources
Anthropic's engineering account of building its own Research feature reports, from internal evaluations, that "multi-agent research systems excel especially for breadth-first queries that involve pursuing multiple independent directions simultaneously". The same post states that a system with Claude Opus 4 as the lead agent and Claude Sonnet 4 subagents outperformed single-agent Claude Opus 4 by 90.2% on its internal research eval.
From the sources
The same post is explicit about where this does not fit: "some domains that require all agents to share the same context or involve many dependencies between agents are not a good fit for multi-agent systems today". As a concrete instance, it notes that most coding tasks involve fewer truly parallelizable tasks than research does.
From the sources
Anthropic's other post, "Building effective agents", gives as its governing principle "finding the simplest solution possible, and only increasing complexity when needed", and says this might mean not building agentic systems at all. It also states that "Agentic systems often trade latency and cost for better task performance", and that you should consider when that trade makes sense.
A&A perspective
Put together, these two replace the reader's question. It is no longer "split or improve" but "is the question in this engagement a breadth-first sweep, or the tracing of a single chain?" A breadth-first sweep is worth splitting. A single chain does not get faster when split. And that distinction can be made from the wording of the request at intake, without touching a line of implementation. What follows is how to make it, and what to write once you have decided to split.
Independence fits in one line: does it need the previous answer?

A&A perspective
The independence test is simple. Take each decomposed direction in turn and ask: to start investigating this direction, do I need another direction's answer? If yes, it is not independent. If no, it is. No technical background is required and no implementation needs to be inspected. All it costs is the effort of splitting the request into directions and writing that one line per direction.
Hypothetical example
Here are two hypothetical requests, placed side by side. Neither is A&A work. Request 1: "list the domestic commercial refrigeration manufacturers that provide their own remote monitoring service". Split into three directions - large manufacturers, small ones, and industry association rosters - no direction waits on another. Request 2: "trace, from public filings, who actually holds voting control of a given holding company". Here the second layer cannot be investigated until the first layer's shareholders are established, and the third layer sits further along still. Both arrive under the same word, "research", but only Request 1 gains from division of labour. Adding agents to Request 2 means either they all wait on the same first layer, or they each set off on an assumption that has not been established.
From the sources
Anthropic's post explains why subagents help by way of the claim that "the essence of search is compression". Subagents, it says, operate in parallel with their own context windows, explore different aspects of the question simultaneously, and then condense the most important tokens for the lead research agent. It adds that each subagent's separate tools, prompts and exploration trajectories are themselves a separation of concerns, which reduces path dependency and enables thorough, independent investigations.
A&A perspective
The practical reading is that the gain comes from separating context, not from the number of agents. Having every agent read the same set of documents and then dividing the work does not separate context. That arrangement carries the same information as one agent reading in sequence, with only the bill increased. If you split, the documents each direction goes to must also differ. If the documents do not differ, there is no reason to split.
| Question shape, identified at intake | Direction of division | Decide first |
|---|---|---|
| Sweeping out several mutually unrelated aspects of the same subject | Split (divide the work) | The four fields per direction, and boundaries that prevent overlap |
| Tracing one chain in order; the next step needs the previous answer | Do not split | Improve the single prompt and tools; resume from interruption |
| Everyone must keep sharing the same current context | Do not split | One place that keeps the context single |
| Raising confidence by applying several views to the same question | Split (double-check) | An output format that lists rather than merges, and the adoption vote count |
| Merely confirming one fact | Do not split | A ceiling on tool calls |
| Showing which conclusion corresponds to which document | Do not split; make it a separate pass | A mapping pass placed after the research loop |
Splitting raises cost by an order of magnitude: how to use the 15x figure
From the sources
The post is candid about cost: "There is a downside: in practice, these architectures burn through tokens fast". In the company's data, agents typically use about 4x more tokens than chat interactions, and "multi-agent systems use about 15x more tokens than chats". It follows with the gate: "For economic viability, multi-agent systems require tasks where the value of the task is high enough to pay for the increased performance".
From the sources
There are also figures on where the performance comes from. On the BrowseComp evaluation, three factors explained 95% of the performance variance, and of those, "token usage by itself explains 80% of the variance". The remaining two were the number of tool calls and the model choice. The post presents this as validating an architecture that distributes work across agents with separate context windows to add capacity for parallel reasoning.
A&A perspective
These numbers need handling with care. Both the 15x and the 90.2% are Anthropic internal evaluations and self-reported data, with no population, period or task mix stated. There is no basis for expecting the same multiples in a small Japanese research practice, and this article does not estimate them. What is usable is not the multiple itself but the fact that cost sits on the order-of-magnitude side. Change the architecture to a split while leaving the price per engagement untouched, and gross margin falls with certainty. So choosing to split always arrives attached to a second decision: move either the price or the scope.
A&A perspective
The order is this. For an engagement you have decided to split, carry four items in the estimate breakdown: the number of directions, the ceiling on inference spend per direction, the cost of the lead role integrating the results, and the cost of the pass that maps conclusions back to sources. If that total does not fit inside the price, the choice is to not split or to narrow the scope. "It gets faster" is not a reason. Whether the contract reflects that speed in either the price or the number of engagements you can accept is the reason.
Once you decide to split, write four fields per direction
From the sources
On what the lead agent must hand to a subagent, the post names four things: "Each subagent needs an objective, an output format, guidance on the tools and sources to use, and clear task boundaries". Without detailed task descriptions, it says, agents duplicate work, leave gaps, or fail to find necessary information.
From the sources
The failure is recorded concretely. Early on, the lead agent was allowed to give short instructions such as "research the semiconductor shortage", and those instructions were often vague enough that subagents misinterpreted the task or performed the exact same searches as other agents. As one instance, "one subagent explored the 2021 automotive chip crisis while 2 others duplicated work investigating current 2025 supply chains" - which the post describes as being without an effective division of labor.
A&A perspective
This transfers directly to delivery work. For each direction you decided to split, write a four-line delegation card. The objective is the one sentence this direction must settle. The output format goes as far as naming the table's columns. The sources name both what to consult and what not to. The boundary is the territory another direction owns, so this one does not touch it. The fourth line does the most work. Duplication presents as a capability problem, but its cause is a boundary that was never written, and an agent does not infer and respect a boundary that is absent.
Hypothetical example
In the hypothetical, splitting Request 1 into three directions looks like this. Direction 1: "for the five largest commercial refrigeration manufacturers, settle from their product pages whether a remote monitoring service is offered in-house. Output five columns: company, product, service name, URL, access date. Boundary: do not handle smaller manufacturers." Direction 2 handles smaller manufacturers in the same five columns, with the five largest excluded by its boundary. Direction 3 works from industry association member rosters and trade show exhibitor lists, adding only company names that Directions 1 and 2 could not reach, with a boundary forbidding re-investigation of companies already covered. This example exists to show the format; it is not a real research result and not an A&A deliverable.
Match the number of agents to the question's complexity
From the sources
The counts are published too. Because agents struggle to judge appropriate effort for a task, the post says, scaling rules were embedded in the prompts. "Simple fact-finding requires just 1 agent with 3-10 tool calls"; direct comparisons might need 2-4 subagents with 10-15 calls each; complex research might use more than 10 subagents with clearly divided responsibilities. The post also records an early failure mode of "spawning 50 subagents for simple queries".
From the sources
"Building effective agents" separates parallelization into two variations. Sectioning is "Breaking a task into independent subtasks run in parallel"; voting is running the same task multiple times to get diverse outputs. And "Parallelization is effective when the divided subtasks can be parallelized for speed", or when multiple perspectives or attempts are needed for higher confidence results.
A&A perspective
In delivery language, sectioning is dividing the work and voting is double-checking. The two are paid for different reasons. Dividing the work is money spent on the deadline; double-checking is money spent against the loss from being wrong. So put them on separate lines in the estimate. Collapsed into the single phrase "we research it with multiple agents", the customer cannot tell whether they bought speed or accuracy, and neither party knows what a discount request is even about. If you include double-checking, decide in advance which conclusions need how many agreeing outputs to be adopted.
A&A perspective
And adding agents to a simple question is precisely the direction the post records as an early failure. A single fact check does not need division of labour. Looking at your list of live engagements and colour-coding which are fact checks and which are breadth-first sweeps settles most of the argument about counts before it starts. If the colour-coding cannot be done, that is not an architecture problem; it is a signal that requests are being accepted too vaguely.
Hiring a person and splitting the work are not on the same axis
A&A perspective
From here on we are outside the sources. Neither referenced post discusses hiring. What follows is A&A's reading. Adding one subagent and adding one person do not line up as interchangeable units. The reason is not whether there are enough hands to investigate, but who writes the design of the division of labour.
From the sources
The post states that multi-agent systems bring a rapid growth in coordination complexity, and that since each agent is steered by a prompt, prompt engineering was the primary lever for improving behaviour. It also states that these systems have emergent behaviours which arise without specific programming, and that small changes to the lead agent can unpredictably change how subagents behave. Success, it says, requires understanding interaction patterns and not just individual agent behaviour, and the best prompts are not strict instructions but "frameworks for collaboration that define the division of labor, problem-solving approaches, and effort budgets".
A&A perspective
So choosing to split is not "more hands to investigate" but "a new job of designing and maintaining the division of labour". How to cut the directions, how to write the boundaries, how to judge the integration, how to chase failures that are hard to predict. That job does not shrink as the agent count rises; it grows. If that is where a one-person business is jammed, what needs adding is not a subagent but a person who can decompose questions and write boundaries. Conversely, if the decomposition is already done and only the hands are short, trying division of labour before hiring is the reasonable order. The deciding input is not a capability comparison but the single question of whether the jam is in decomposition or in execution.
From the sources
The maintenance cost is covered as well. Because agents make dynamic decisions and are "non-deterministic between runs, even with identical prompts", debugging is harder. Users would report agents not finding obvious information without the reason being visible, and only adding full production tracing made it possible to diagnose failures systematically. The post also states that agents are stateful and errors compound, so the team built systems that can resume from where an error occurred, combined with deterministic safeguards such as retry logic and regular checkpoints.
A&A perspective
For a one-person business this means that introducing division of labour brings a requirement for observation tooling. If you cannot see afterwards what each direction searched and where it came up empty, you cannot identify what to fix when quality drops. In delivery work where deliverable quality is promised, that is not an optional luxury but a precondition. If you are going to try splitting, first build the place where each direction's search record and failures are kept. If you cannot build it, then "do not split on this engagement" is the correct conclusion.
Where this judgement does not hold, and the next step
A&A perspective
Four limits. First, every figure quoted is an Anthropic internal evaluation or self-reported data, not an independent audit. Neither the 90.2% nor the 15x shares its evaluation target, population or task mix with a small Japanese research practice; this article uses them only as a sense of order of magnitude and never as a cost-effectiveness estimate. Second, Anthropic supplies the models, a position from which a multi-agent architecture looks favourable. Third, "Building effective agents" is dated December 2024 and carries its own note that "the tooling landscape described in this post has changed since December 2024". Its design principles remain readable, but any specific tooling claim needs rechecking against the present state. Fourth, both posts address developers building products, not the contracts, acceptance criteria and liability of a practice that delivers research to customers. All of that part is A&A's interpretation.
A&A perspective
There are also cases where the thesis fails. Even when the question decomposes breadth-first, if what the customer wants is one decisive document rather than coverage, the gain from splitting does not become deliverable value, because what is being paid for is arriving at the one central document rather than a wide table assembled from ten directions. The other case is when the material sits inside one limited set - a customer's internal documents, or a single database. Splitting directions does not widen the search space there; it only re-reads the same material separately. That case is close in shape to what the source itself excludes as domains requiring all agents to share the same context.
A&A perspective
The next step is not implementation. Write down the three research engagements you currently hold, decompose each question into directions, and write one line per direction on whether it needs another direction's answer. If not one engagement has two or more independent directions, division of labour needs no consideration right now. If at least one does, write the four-line delegation cards for that engagement alone, and add review time for the number of directions to the estimate. How to estimate that post-delivery review time itself is handled in "Before taking on more AI delivery work, estimate where human review jams". For where this decision sits in the flow from winning customers through delivery to retention, see "AI-native GTM: a practical guide for solo founders and small teams". If the decomposition is hard on every single engagement, that is not an architecture problem but a requirement-definition problem at intake, and it is the kind of question suited to a direct consultation.
Split, improve, or hire. Trying to settle that three-way choice first produces no answer. It settles after you have written down whether the question you accepted decomposes into independent directions. Write the four fields per direction only for the engagements that decompose, and put the cost and review time for that number of directions into the estimate. For engagements that do not decompose, improve the single prompt and its tools rather than the agent count. The figures quoted are another company's internal evaluations - not our own measurements, and not a forecast of your cost-effectiveness.
Sources & editorial note
Primary pages read for this article. Publication dates below belong to the sources; access dates record our research.
- How we built our multi-agent research system
Anthropic · Jun 13, 2025 (as shown on the page)
Accessed 2026-09-30 - Building effective agents
Anthropic · Dec 19, 2024 (page date); page notes tooling has changed since
Accessed 2026-09-30
AI-assisted editorial production
A&A uses AI for research, writing, translation and editorial checks. Source facts, our analysis and hypothetical examples are labeled separately.
Editorial check: 2026-09-30