Most builders learn this from an invoice. An agent re-sends its entire conversation on every single turn, so cost grows with the square of the turn count, not in a straight line.

You built an agent. It works. It solves the task in about thirty tool calls and produces a conversation that, if you copied the whole transcript into a text file, would be maybe fifty thousand tokens long. Then the bill arrives and you were charged for something closer to three-quarters of a million input tokens. Nothing is broken. Nothing was double-charged. This is simply how the agent loop works, and almost nobody explains it before you pay for it.

This article covers the mechanism: why an agent's cost curve bends upward, which parts of the request are re-billed on every turn, and where the spend actually concentrates. The reasoning is provider-neutral; the worked numbers use Claude's published rates because they are documented precisely enough to check. Every technique named here has a tradeoff, and each one is stated.

All prices, limits, and token counts in this article are as of August 2026 and are quoted from primary provider documentation. Rates change frequently — verify current pricing before making a budget decision on these figures.

The mechanism: LLM APIs are stateless, agents are not

The Messages API — and every equivalent chat-completions endpoint — is stateless. It has no memory of your previous request. Every time you call it, you send the whole conversation: the system prompt, the tool definitions, every user message, every assistant reply, and every tool result produced so far. The model reads all of it, then produces one more assistant turn.

An agent is a loop wrapped around that stateless call. Send request. Model replies with a tool call. Your code executes the tool. Append the result. Send the whole thing again. Repeat until the model stops asking for tools. Each iteration is a separate billable request, and each one carries everything that came before it. That is the entire mechanic — not subtle, not a bug, and the arithmetic it produces surprises almost everyone.

A worked example: one task, thirty turns

Take a modest agent. Its fixed prefix — system prompt plus tool schemas — is 3,000 tokens, and the user's task adds 500. On each turn the model produces about 300 output tokens (some thinking, a tool call) and the tool returns about 1,200 tokens of result, so each completed turn permanently adds roughly 1,500 tokens. The input billed on turn n is therefore 3,500 + 1,500 x (n - 1), and the total across n turns is the sum of that series.

prefix       = 3,000 (system prompt + tool schemas)
task         =   500 (user message)
per turn     = 1,500 (300 model output + 1,200 tool result)
 
input on turn n     = 3,500 + 1,500 x (n - 1)
total input, n turns = n x 3,500 + 1,500 x n x (n - 1) / 2
After turn Input billed this turn Cumulative input billed Live conversation size Billed / live
1 3,500 3,500 5,000 0.7x
5 9,500 32,500 11,000 3.0x
10 17,000 102,500 18,500 5.5x
20 32,000 355,000 33,500 10.6x
30 47,000 757,500 48,500 15.6x

Read the last two columns together. At the end of a thirty-turn task the conversation itself is 48,500 tokens. You were billed for 757,500 input tokens to produce it — roughly sixteen times the size of the artifact. Nothing was wasted; every one of those tokens was genuinely read by the model. It just read the same early tool results thirty times.

This is why cost is quadratic, not linear

Ten turns cost 102,500 input tokens. Thirty turns cost 757,500. Tripling the turn count multiplied cost by 7.4x, not 3x. The n(n-1)/2 term dominates once the fixed prefix stops mattering, which means input cost scales with roughly the square of the turn count.

Priced at Claude Sonnet 5 rates as of August 2026 ($2 per million input tokens, $10 per million output), the ten-turn version costs about $0.235 and the thirty-turn version about $1.61 — for the same user request, just solved less directly. That ratio, not the absolute figure, is the thing worth internalising. Wandering agents are not slightly more expensive. They are quadratically more expensive.

The single most useful cost intuition for agent builders: every additional turn is charged the price of the entire conversation so far. A turn that goes nowhere is not cheap — it permanently raises the price of every turn after it.

Where the spend actually concentrates

1. Tool results, not prompts

In almost every real agent, tool results are the largest single contributor to context growth. A prompt you wrote is a few hundred tokens and you can see it; a tool result is whatever the tool returned and you usually cannot. Anthropic's documentation gives useful scale markers for fetched content: an average web page of 10 kB is about 2,500 tokens, a large documentation page of 100 kB about 25,000, and a research-paper PDF of 500 kB about 125,000.

One careless fetch can therefore add more to your context than the entire rest of the conversation — and you pay for it on every subsequent turn. The web fetch tool exposes a max_content_tokens parameter precisely for this. Shell tools are worse, because stdout is unbounded: a test suite that prints every passing assertion, an unfiltered log tail, or a directory listing of node_modules all land in context at full size.

Tradeoff: truncating tool output saves tokens on every future turn but risks cutting the one line the model needed. Truncate with a filter you control — grep for failures, head the output, return structured summaries — rather than a blind character cap.

2. Tool and function schemas, billed on every call

Your tool definitions are part of the input on every single request. So is a provider-inserted system prompt that teaches the model how to use tools at all. Anthropic publishes these overheads exactly, which makes them checkable — figures below are as of August 2026.

Model Tool-use system prompt (tool_choice auto or none) Tool-use system prompt (any or tool)
claude-opus-5 290 tokens 410 tokens
claude-sonnet-5 354 tokens 474 tokens
claude-haiku-4-5 496 tokens 588 tokens

Individual tools add more on top. On claude-opus-5 the bash tool definition adds about 325 input tokens; the text editor tool text_editor_20250429 adds about 700. The heavier toolsets are where this stops being rounding error: declaring computer_toolset_20260801 with its default members adds roughly 4,520 input tokens on claude-opus-5, and browser_toolset_20260801 adds roughly 6,610. Those are paid on turn one and on turn thirty alike.

Your own tool schemas behave the same way. Twenty tools with verbose descriptions and deeply nested JSON schemas can easily reach several thousand tokens of pure overhead re-billed per turn — over a thirty-turn task, a 4,000-token tool block costs 120,000 input tokens on its own.

Tradeoff: tool descriptions are how the model decides which tool to call, so cutting them too far buys back the savings in wrong tool calls, which are themselves extra turns. Prefer removing whole unused tools over shortening the ones you keep. Deferred tool loading — only tool names in context until the model asks for a schema — is the structural fix where your platform supports it.

3. Reasoning and thinking tokens are output tokens

Extended or adaptive thinking is billed as output. That matters because output is the expensive side of the ledger: as of August 2026, claude-opus-5 is $5 per million input and $25 per million output, claude-sonnet-5 is $2 and $10, and claude-haiku-4-5 is $1 and $5. Across the current lineup, output costs five times input.

Claude Code's documentation is blunt about the scale: the default thinking budget can be tens of thousands of tokens per request. On an agentic loop that reasoning happens every turn, not once. The current control is the effort parameter rather than a fixed budget.

Tradeoff: lowering effort reduces cost per turn but raises the risk of a wrong approach — and a wrong approach costs turns, which are quadratic. Cheap reasoning that produces ten extra turns of flailing is a net loss.

4. Retries, error loops, and the invisible multiplier

Every failed tool call is billed twice over. You pay for the turn that produced the bad call, you pay for the error message that comes back as a tool result, and then you pay for both of them again on every remaining turn because they are now permanently in the conversation. A malformed argument that triggers three correction rounds does not add three turns of cost — it adds three turns near the expensive end of the curve and inflates the baseline for everything after.

The pathological version is the loop where the model retries the same failing action with cosmetic variations — the strongest argument for a hard iteration cap in your own loop code, independent of anything the model does. Tradeoff: a hard cap turns runaway spend into a failed task. Usually the right trade, but the failure needs to be visible and resumable rather than silent.

5. Server-side tools with their own meters

Some tools cost money beyond tokens. Anthropic's web search is $10 per 1,000 searches as of August 2026 on top of standard token costs — and those results count as input tokens both in the turn that ran them and in every subsequent turn. Web fetch has no per-call charge but you pay full token price for whatever it drags in. Code execution is free alongside web search or web fetch; used alone it is billed by execution time, with 1,550 free container-hours per organisation per month and $0.05 per hour per container beyond that. Individually small, easy to omit from a cost model entirely — and not small at all in a research agent running eight searches per task.

What prompt caching does and does not fix

Prompt caching is the direct countermeasure to re-sent context, and it is genuinely effective. Anthropic's published multipliers as of August 2026: a five-minute cache write costs 1.25x the base input rate, a one-hour write costs 2x, and a cache read costs 0.1x — so caching pays for itself after a single read on the five-minute TTL, or two reads on the one-hour TTL. Applied to the thirty-turn example, the bulk of that 757,500 input tokens is history being re-read; if most of it lands as cache reads, the effective input bill drops by something in the region of four to five times, with no change to what the agent actually does.

But caching does not repeal the mechanism. Three things limit it:

  • It is a prefix match, and the hierarchy is tools, then system, then messages. Change a tool definition and you invalidate the tools, system, and message caches simultaneously. A tool list that varies by request — assembled from a dict, or including a timestamp — silently caches nothing.
  • There is a minimum cacheable prefix. As of August 2026 it is 1,024 tokens on claude-opus-5 and claude-sonnet-5, and 4,096 on claude-haiku-4-5. Shorter prefixes fail to cache with no error raised.
  • Caches expire. The default TTL is five minutes. An agent that pauses for user input, or a session resumed after lunch, starts cold and reprocesses everything at full price.

Verify it rather than assume it: check cache_read_input_tokens in the response usage. If it is zero across requests that should share a prefix, something is invalidating your cache and you are paying full rate for the whole conversation on every turn.

Tradeoff: the one-hour TTL costs 2x on write instead of 1.25x, so it only wins if you are confident of at least two reads. Paying 2x for a cache that never gets read is a straightforward loss.

If you are on a subscription plan rather than an API bill

On a flat monthly plan you do not see dollars, but you see the same mechanism as a usage limit arriving sooner than expected. Claude Code's documentation says it directly: it sends your full conversation with every request, and each time Claude uses tools it sends another request carrying that batch of tool results — so a one-line question in a session that has been open all day still draws usage for the whole conversation.

The industry has converged on token metering regardless of the interface. Cursor bills per token across input, cache read, cache write and output rather than per request, drawing from monthly usage pools. GitHub Copilot moved to token-based billing denominated in AI Credits at 1 credit = $0.01, with per-million-token model rates — though code completions and next edit suggestions remain unbilled and unlimited on paid plans. The abstraction differs; the underlying quadratic does not.

Where to look first

In rough order of impact per unit of effort:

  1. Measure before optimising. Sum output_tokens across every request in your loop and add the token size of the tool results you append between requests. Most teams discover their spend is concentrated somewhere they did not expect.
  2. Find the biggest tool result. Log the token size of every tool return. There is usually one tool that dwarfs the rest, and capping it is the single highest-leverage change available.
  3. Confirm caching is actually hitting. A zero cache-read count is a five-minute fix worth more than a week of prompt trimming.
  4. Count the turns. If a task that should take eight turns takes twenty-five, the fix is a better task specification, not a cheaper model.
  5. Cap iterations in your own loop code. Not as an optimisation — as a circuit breaker.

Model choice comes after all of these. Moving from claude-opus-5 to claude-sonnet-5 halves input cost and cuts output cost by 60%, which is real. But it does nothing about a loop that reads a 40,000-token file on turn three and then carries it for twenty-seven more turns. Fix the context first; the model tier is a multiplier applied to whatever you have left.

  • Token hygiene for coding agents — the practical habits that stretch limits in Claude Code, Cursor, Copilot, and Replit.
  • Multi-agent cost control — when subagents pay for themselves and when they multiply the bill.
  • Building your first Claude agent: tool use, the agent loop, and streaming — the loop mechanics this article prices.
  • Replit Agent credit consumption: why costs are unpredictable — the same mechanism seen through a credit-based interface.
  • LiteLLM: one API to route all your LLMs — where to put per-key spend tracking once you know what to track.
  • Langfuse 101 — instrumenting an agent so the token distribution above is observable rather than inferred.
Re-audit reminder: model IDs, per-token rates, cache multipliers, tool-overhead token counts, and server-tool pricing in this article were verified against Anthropic, Cursor, and GitHub primary documentation in August 2026. Treat every figure as a dated snapshot.