Every subagent call is a full LLM invocation with its own context. Whether that saves money or triples it comes down to one question: how much context do your subagents share?

The standard advice about subagents contains a real insight and a dangerous omission. The insight: delegating a verbose task to a subagent keeps its noisy output out of your main context, so you stop paying for it on every subsequent turn. The omission: the subagent still ran. It had its own system prompt, its own tool schemas, its own file reads, and its own multi-turn loop — all separately billed. What you saved in your context, you spent in theirs.

Whether that trade comes out ahead is not a matter of opinion. It is arithmetic, and it hinges on a single variable: how much context the subagents need to share. This article works that arithmetic, shows the two cases where it flips, and covers the controls that keep a fan-out from becoming a runaway. AI Workshack already covers the orchestration patterns themselves — supervisor graphs in LangGraph, crew structures in CrewAI, handoffs in the OpenAI Agents SDK, orchestrator-subagent design with the Claude Messages API. This is about what those patterns cost, not how to build them.

All prices, token counts, and product behaviours here are as of August 2026 and were verified against primary provider documentation. Rates and agent-platform features change frequently — verify before budgeting on these figures.

The cost model in one formula

From the mechanics covered in 'Why AI agents burn tokens': an agent re-sends its whole conversation every turn, so the input billed across an n-turn conversation is approximately:

total input  =  n x prefix  +  growth x n x (n - 1) / 2
 
prefix  = system prompt + tool schemas + any context loaded up front
growth  = tokens permanently added per completed turn
          (model output + tool result)

Two terms, and multi-agent designs move them in opposite directions:

  • The quadratic term shrinks when you split work. Four twelve-turn conversations produce far less n-squared growth than one forty-eight-turn conversation. This is the real saving, and it is substantial.
  • The linear term multiplies. Each agent pays its own prefix on every one of its turns. Add a fifth agent and you add a fifth full copy of the system prompt, tool schemas, and any shared context — re-billed on each of its turns.

So the question is never 'are subagents expensive'. It is whether the prefix you duplicate is smaller than the quadratic growth you avoid.

Worked example: the same task, three ways

The baseline is the thirty-turn single agent from the earlier article: a 3,500-token prefix plus task, 1,500 tokens of growth per turn, 300 output tokens per turn. That comes to 757,500 input tokens and 9,000 output tokens. At claude-sonnet-5 rates as of August 2026 ($2 per million input, $10 per million output) it costs about $1.61.

Case A: disjoint fan-out — subagents win

The work genuinely partitions into four independent slices. An orchestrator spawns four subagents, each of which needs only its own slice. Each subagent carries a 4,000-token prefix (its own system prompt and tools) and runs twelve turns. The orchestrator runs eight turns, and what accumulates in its context is subagent summaries of about 800 tokens each rather than raw tool output.

Component Turns Prefix Input tokens Output tokens
Orchestrator 8 3,500 61,600 3,200
Subagent x 4 12 each 4,000 each 588,000 total 14,400 total
Multi-agent total — — 649,600 17,600
Single agent baseline 30 3,500 757,500 9,000

The multi-agent version costs about $1.48 against the single agent's $1.61 — roughly 8% cheaper, and it ran four slices in parallel. The extra output tokens (each subagent writes its own summary) partly offset the input saving, but the quadratic term did most of the work: four twelve-turn conversations simply generate less re-sent history than one thirty-turn conversation.

Case B: overlapping fan-out — subagents lose

Same structure, one difference: the work does not cleanly partition, so every subagent needs the same 20,000 tokens of shared context — the spec, the schema, the relevant module. Its prefix is now 24,000 instead of 4,000.

Component Turns Prefix Input tokens Output tokens
Orchestrator 8 3,500 61,600 3,200
Subagent x 4 12 each 24,000 each 1,548,000 total 14,400 total
Multi-agent total — — 1,609,600 17,600
Single agent baseline 30 3,500 757,500 9,000

About $3.40 against $1.61 — roughly 2.1x the single-agent cost, for the same task. Nothing went wrong. No agent looped. The 20,000-token shared context was simply paid four times over, on twelve turns each: 1,152,000 tokens of pure duplication.

Case C: add one retry round

The orchestrator judges two of the four results inadequate and re-runs them. That adds two more full subagent runs at 387,000 input tokens each — pushing the total past 2.38 million input tokens, or roughly $4.98. Around 3.1x the single-agent cost, and this is the shape of the multi-agent bills people are surprised by. Not a runaway loop: just a design where subagents share most of their context, with one round of quality control on top.

The break-even test, before you build anything: write down how many tokens of context each subagent genuinely needs that the others do not. If the answer is 'most of it is shared', a single agent with good context discipline will be cheaper and simpler. Fan out on genuinely disjoint work.

What a subagent actually costs to start

In Claude Code, a fresh subagent's initial context contains its own system prompt plus environment details, the delegation prompt written by the parent, CLAUDE.md files from every level of the hierarchy, a git status snapshot, the full content of any preloaded skills, and a roster of sibling agents. It does not see your conversation history, previously invoked skills, or files the parent already read.

That cuts both ways: it is why subagents keep verbose output out of your context, and why they re-read things the parent already knows. If your CLAUDE.md is 1,800 tokens and you spawn five subagents that each run ten turns, that file alone is billed fifty more times. One exception is worth flagging — a fork inherits the entire parent conversation rather than starting fresh, which is the most expensive possible way to spawn a helper: you have duplicated the whole context you were trying to escape.

The heavier tier: independent agent instances

Claude Code distinguishes subagents from agent teams, where each teammate is a separate instance with a fully independent context window. The documented cost difference is stark: agent teams use approximately 7x more tokens than standard sessions when teammates run in plan mode, and usage scales roughly linearly with team size. Anthropic's guidance is to start with 3 to 5 teammates and keep spawn prompts focused, because teammates load CLAUDE.md, MCP servers and skills automatically and everything in the spawn prompt adds to their context from the start.

There is also a caching trap: an in-process teammate's requests fall outside the main conversation's cache TTL bucket, so its cache holds for five minutes by default even on a subscription where the main session gets an hour. A setting raises it to an hour, but one-hour cache writes are billed at 2x base input rather than 1.25x, so that only pays off for a genuinely long-running teammate.

Tradeoff: independent instances buy real parallel exploration and genuine adversarial review — teammates that challenge each other's hypotheses find root causes that a single agent, anchored on its first plausible theory, does not. Worth 7x on a hard debugging session. Not worth 7x on a refactor.

When parallelism is actually worth it

Fan-out earns its cost in a narrow set of shapes:

  • Independent exploration where anchoring is the enemy. Multiple investigators testing competing hypotheses avoid the failure mode where one agent finds a plausible explanation and stops looking. You are buying diversity of search, not throughput.
  • Reading-heavy work with a small output. A subagent that reads twenty files and returns 400 tokens is a clear win: the reads stay in a context you throw away, and only the summary enters the conversation you keep paying for.
  • Genuinely partitioned work. Different files, different services, different review lenses — anywhere the subagents' context requirements barely overlap.
  • Wall-clock latency matters more than token cost. Sometimes the right answer is simply that you are buying time, and the extra tokens are the price. That is a legitimate reason; just name it as the reason.

And where it does not:

  • Sequential work with dependencies. If agent two cannot start until agent one finishes, you have paid the duplication cost and bought no parallelism.
  • Same-file edits. You pay N agents to conflict with each other and then pay again to reconcile.
  • Tasks small enough that coordination overhead exceeds the work. Anthropic's guidance on sizing is direct: too small and coordination overhead exceeds the benefit.
  • Anything where every worker needs most of the shared context. This is Case B above, and it is the most common way multi-agent designs get expensive.

Capping iterations and depth

The failure mode that produces genuinely alarming bills is unbounded recursion: agents spawning agents, or a loop that retries indefinitely. Three layers of control, from most to least reliable.

1. A hard iteration cap in your own loop

If you wrote the agent loop, you own the ceiling. A maximum-iterations counter that raises an error is the only genuinely enforced control, because it lives in your code rather than the model's judgement. The same applies to spawn depth: a counter passed down through each delegation, refusing to spawn below a fixed depth. Some platforms build this in — Claude Code's agent teams do not permit nested teams, so the fan-out is one level deep by construction.

2. An advisory budget the model can see

The Claude API offers task budgets in beta: a token ceiling for the full agentic loop that the model sees as a running countdown and paces itself against, so it finishes gracefully rather than being cut off mid-action. It goes inside output_config with the beta header task-budgets-2026-03-13.

with client.beta.messages.stream(
    model="claude-opus-5",
    max_tokens=128000,
    output_config={
        "effort": "high",
        "task_budget": {"type": "tokens", "total": 64000},
    },
    betas=["task-budgets-2026-03-13"],
    messages=[{"role": "user", "content": "..."}],
    tools=[...],
) as stream:
    response = stream.get_final_message()

Three things to know before reaching for it. It is advisory, not enforced — Claude may exceed the budget when interrupting an action would be more disruptive than finishing it, so max_tokens remains your hard per-request ceiling. Model support is uneven: as of August 2026 it is available on claude-opus-5 but not on claude-sonnet-5 or claude-haiku-4-5, and it is not supported on Claude Code. And the documented minimum total is 20,000 tokens.

A task budget that is too small for the work causes refusal-like behaviour: Claude may decline the task, scope it down aggressively, or stop early with a partial result rather than start work it cannot finish. Size budgets against a measured distribution of your real task lengths — Anthropic's guidance is to start from the p99 of per-task token spend — rather than picking a round number.

3. Platform-level spend caps

Workspace spend limits, per-key caps on an LLM gateway, or organisation budget controls in your cloud provider. These do not prevent a bad run; they bound the blast radius. Treat them as the backstop, not the control.

Cheaper models for narrow subagents

Subagent work is often the best candidate for a cheaper tier, because a well-scoped subagent task is narrow, mechanical, and verifiable — run the tests and report failures, fetch these docs and summarise, scan for this pattern. Across the current Claude lineup as of August 2026 the input tiers run roughly 5:2:1 (claude-opus-5 at $5 per million input and $25 output, claude-sonnet-5 at $2 and $10, claude-haiku-4-5 at $1 and $5), so routing narrow work down a tier is a direct multiplier on the largest part of a fan-out bill.

In Claude Code the model is set per subagent definition, with a documented resolution order: the CLAUDE_CODE_SUBAGENT_MODEL environment variable, then a per-invocation model parameter, then the subagent definition's frontmatter, then the main conversation's model. Anthropic's guidance for agent teams is to use Sonnet for teammates, as a balance of capability and cost for coordination tasks.

Tradeoff: this is where cheap gets expensive. A subagent that returns a wrong or incomplete summary does not just waste its own run — it feeds bad input to the orchestrator, which spends more turns discovering the problem and then re-runs the subagent. Because each subagent run is a full conversation, a retry is not a small correction; it is another whole invocation. Downgrade subagents whose output the orchestrator can cheaply verify. Keep the stronger model where a wrong answer is expensive to detect.

Measuring cost per task

None of the above is decidable without measurement, and aggregate spend will not tell you. You need cost attributed to a task, with subagent runs rolled up into it.

  1. Assign a task ID at the top of the orchestration and propagate it to every subagent invocation. Without this, subagent spend appears as unattributed background noise.
  2. Record per-request usage everywhere, not just at the top level: input tokens, output tokens, cache creation tokens, and cache read tokens as four separate figures. Collapsing them into one number hides the single largest cost signal, because a cache read costs a tenth of base input.
  3. Roll up to a cost per completed task, and track its distribution rather than its mean. Multi-agent cost distributions have long tails; the p95 is what causes incidents.
  4. Track the failure rate alongside it. Cost per attempted task and cost per successfully completed task diverge sharply once retries enter the picture, and only the second number is meaningful.
  5. Run the single-agent control. Before committing to a fan-out, measure the same task solved by one agent. Surprisingly often the fan-out is not faster and not cheaper, only more complicated.

If your platform surfaces attribution natively, use it. Claude Code's /usage breakdown on paid plans attributes recent usage to skills, subagents, plugins and individual MCP servers as a percentage of the total, which is the fastest way to discover that delegation is consuming more of your allowance than the main conversation. For API integrations, an observability layer or a gateway that tracks spend per key gives you the same picture across providers.

A decision checklist

Question If yes If no
Do the subtasks need mostly disjoint context? Fan out — the quadratic saving is real Stay single-agent; you will duplicate the prefix N times
Is the subagent's output much smaller than its input? Delegate — this is the classic win Delegation buys little; the summary is nearly as big as the work
Can the orchestrator cheaply verify each result? A cheaper model per subagent is safe Keep the stronger model; retries cost a full run
Is the work sequential or dependency-heavy? Single agent or a simple chain Parallel fan-out is viable
Is wall-clock latency the actual constraint? Pay the token premium deliberately Optimise for tokens instead
Do you have per-task cost attribution in place? Proceed and measure Build the measurement first — you cannot tune what you cannot see

The summary is short. Subagents do not save money by existing. They save money when they let you avoid carrying context you do not need, and they cost money when they force you to carry the same context several times. Everything else — model tiering, iteration caps, budgets — is a multiplier applied on top of that structural decision.

  • Why AI agents burn tokens — the per-turn mechanic this article's arithmetic is built on.
  • Token hygiene for coding agents — the single-agent discipline to try before reaching for a fan-out.
  • Multi-agent systems with Claude: orchestrators, subagents, and when to split — the design patterns priced here.
  • LangGraph multi-agent supervisor: building agents that delegate.
  • OpenAI Agents SDK: building multi-agent systems with handoffs.
  • Running CrewAI in production: deployment, cost control and rate-limit handling.
  • Langfuse 101 and LiteLLM — the instrumentation and gateway layers that make per-task cost attribution practical.
Verification note: subagent context composition, agent-team token multiples and caching behaviour, subagent model resolution order, and task-budget syntax, model support and minimums were checked against Anthropic primary documentation in August 2026. Per-token rates are dated snapshots. The worked examples are illustrative models built on those documented mechanics, not measurements of a specific production system — run the arithmetic with your own prefix and growth figures.