Your history grows and gets trimmed. Your system prompt and tool schemas are re-sent in full, unchanged, on every single request — forever.
Conversation history is the token cost everybody talks about. It is also the one that self-corrects: you trim it, you summarise it, it goes away. The cost nobody plots is the one that never goes away — the system prompt and the tool schemas that are re-serialised and re-sent, byte for byte identical, on every request your application will ever make.
This is a flat tax. It does not scale with conversation length, it does not respond to trimming, and it is invisible in a per-request token count because it looks like a constant. Multiply that constant by your request volume and it is frequently the largest single line on the bill. It is also, once you know the two techniques in this article, one of the easiest things to fix.
All token counts, pricing multipliers, cache minimums and API parameters in this article were verified against Anthropic and OpenAI documentation as of August 2026. Per-model overheads and cache minimums change with every model generation — re-measure against your own models rather than trusting any figure here, and verify current pricing before building a business case on it.The mechanism: fixed overhead times request count
There is no clever framing here. If your system prompt and tool definitions total 12,000 tokens and you serve a million requests a month, you have paid for twelve billion input tokens before a single user has typed anything. Whether that matters depends entirely on one variable — request volume — and on one technique, prompt caching, which we will get to.
What makes this hard to notice is that the overhead is not itemised anywhere. Your usage object reports input_tokens as a single number. The system prompt, the tool schemas, the provider's own injected tool-use preamble and the actual user message all arrive as one figure. To manage it you have to decompose it yourself.
Measuring it: the delta method
The token counting endpoint is stateless, so you can isolate any component by counting with and without it and subtracting. Four counts give you a full breakdown, and this is worth wiring into CI as a regression test — prompt bloat accretes one helpful paragraph at a time, and nobody notices the day it doubles.
import anthropic
client = anthropic.Anthropic()
MODEL = "claude-sonnet-5"
PROBE = [{"role": "user", "content": "."}] # near-zero user content
def count(**kwargs) -> int:
return client.messages.count_tokens(
model=MODEL, messages=PROBE, **kwargs
).input_tokens
baseline = count() # probe only
with_system = count(system=SYSTEM_PROMPT)
with_tools = count(tools=TOOLS)
with_both = count(system=SYSTEM_PROMPT, tools=TOOLS)
print(f"system prompt: {with_system - baseline:>6,}")
print(f"tools + tool preamble:{with_tools - baseline:>6,}")
print(f"total fixed overhead: {with_both - baseline:>6,}")
# Per-tool attribution: which schema is actually the problem?
for tool in TOOLS:
others = [t for t in TOOLS if t["name"] != tool["name"]]
print(f" {tool['name']:<28} {count(tools=TOOLS) - count(tools=others):>6,}")That last loop is the one that changes minds. Teams routinely discover that a single tool — usually one with a deeply nested input schema, or one auto-generated from a database model — accounts for a third of their tool budget.
The overhead you did not write
Passing any tools at all makes the provider inject its own tool-use system prompt, which you never see and always pay for. It is model-specific, and it is not stable across generations — which means a model upgrade can change your fixed cost without you changing a line of code.
| Model | tool_choice auto / none | tool_choice any / tool |
|---|---|---|
| Claude Opus 4.8 | 290 tokens | 410 tokens |
| Claude Sonnet 5 | 354 tokens | 474 tokens |
| Claude Haiku 4.5 | 496 tokens | 588 tokens |
Anthropic-provided tools add their own definitions on top. The bash tool adds about 325 tokens on the current Opus generation; the text editor tool about 700. The larger toolsets are in a different league entirely: declaring the computer use toolset with its default members adds roughly 4,500 input tokens per request, and the browser use toolset roughly 6,600. Those are not typos. If you have declared a browser toolset for a feature that fires on one request in fifty, you are paying that overhead on the other forty-nine.
If you pass no tools at all and set tool_choice to none, the injected preamble costs zero additional tokens. The cheapest tool is the one you did not declare on this request.Why tool schemas are usually the bigger offender
Almost every team assumes the system prompt is the problem, because the system prompt is the thing they wrote and can see. In practice the tool schemas are usually larger, for four structural reasons.
- JSON Schema is verbose by construction. Every property carries a type, a description and often an enum. The braces, quotes and repeated keys tokenize poorly compared to prose — you are paying for punctuation.
- Descriptions are duplicated at every level. A well-written tool has a description on the tool, on each property, and on each enum member. That is good practice for accuracy and expensive for tokens.
- Auto-generation has no editor. Schemas produced from Pydantic models, OpenAPI specs or database introspection include every field the source object has, including the ones the model has no business setting.
- MCP servers multiply it. Connecting a server means importing its entire tool surface, not the two tools you wanted. Anthropic's own documentation puts a typical five-server setup — GitHub, Slack, Sentry, Grafana and Splunk — at roughly 55,000 tokens of definitions before the model does any work at all.
Fifty-five thousand tokens of fixed overhead, on every request, in a product where most requests need three tools. That is the shape of the problem.
How many tools is too many?
There is a quality answer and a cost answer, and helpfully they point the same direction. On the quality side, Anthropic's documentation states that the model's ability to select the right tool degrades once you exceed roughly 30 to 50 available tools. On the cost side, the guidance for reaching for on-demand tool loading is any of: ten or more tools, tool definitions exceeding 10,000 tokens, aggregation of multiple MCP servers, or a tool library that grows over time. Below ten tools with small definitions — under about 100 tokens total — plain tool calling is the right answer and this is not your problem.
| Toolset size | Cost posture | Accuracy posture | Recommended approach |
|---|---|---|---|
| Under 10, small schemas | Negligible fixed cost | Reliable selection | Plain tool calling. Do nothing. |
| 10–30 | Worth measuring; often >10k tokens | Generally reliable | Cache the tool block; audit the largest schemas |
| 30–50 | Material fixed cost | Selection accuracy starts degrading | Trim schemas hard, or move to deferred loading |
| 50+ / multi-MCP | Dominates the request | Selection degrades noticeably | Deferred loading (tool search) is the intended answer |
Trimming: what actually shrinks a schema
- Delete parameters the model should never set. Anything derived from the session — user IDs, tenant IDs, auth context, timestamps — belongs in your handler, not in the schema. This is a security improvement as much as a cost one.
- Collapse near-duplicate tools. Four tools that differ by one enum value should be one tool with an enum. You pay for one description and one schema instead of four.
- Shorten descriptions to the decision, not the documentation. A tool description exists to help the model choose between tools. It is not an API reference for humans. Cut usage notes, examples of output, and error semantics — the model does not need them to decide.
- Prune auto-generated schemas rather than shipping them raw. Hand-write the schema for your five most-used tools; generation is a starting point, not an output.
- Drop tools that are not reachable on this request. If admin tools only apply to admin sessions, do not send them to everyone — but read the caching caveat below before you make this dynamic.
- Move the long instructions out of the system prompt and into a tool or file the model reads on demand. Instructions that apply to one task in twenty should be loaded for that one task.
Point five has a trap. Varying the tool list per user or per request is excellent for raw token count and terrible for caching: tool definitions render at position zero, so any variation gives every variant its own cache entry and none of them accumulate reads. If you cache, a slightly larger frozen tool list usually beats a smaller dynamic one. Deferred loading, below, is how you get both.Deferred loading: send the schemas without paying for them
The tool search tool separates two things that used to be the same thing: which tools the API knows about, and which tool definitions occupy the context window. You still send every definition on every request — the server needs them to run the search — but the ones marked defer_loading are excluded from the prefix until the model actually searches for and finds them. Anthropic reports this typically cuts definition overhead by more than 85 percent, loading only the three to five tools a given request needs.
import anthropic
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-opus-5",
max_tokens=2048,
messages=[{"role": "user", "content": "What is the weather in San Francisco?"}],
tools=[
# The search tool itself must NEVER be deferred.
{"type": "tool_search_tool_regex_20251119",
"name": "tool_search_tool_regex"},
# Keep your 3-5 most-used tools loaded up front.
{"name": "get_weather",
"description": "Get the weather at a specific location",
"input_schema": {
"type": "object",
"properties": {"location": {"type": "string"}},
"required": ["location"],
}},
# Everything else: present in the request, absent from context
# until Claude searches for it.
{"name": "search_files",
"description": "Search through files in the workspace",
"input_schema": {
"type": "object",
"properties": {"query": {"type": "string"}},
"required": ["query"],
},
"defer_loading": True},
],
)Two constraints matter in practice. At least one tool must stay non-deferred or the request is rejected with a 400 — normally the search tool itself satisfies this. And a tool carrying defer_loading cannot also carry cache_control; put your cache breakpoint on a non-deferred tool.
The tradeoffs are real. Discovery costs a round trip: the model searches, the API returns references, and only then can it call the tool — so a request that needs a deferred tool is slower than one that does not. Discoverability becomes a design problem: a tool whose description does not contain the words the model would search for is effectively invisible, which is a new class of bug. And tool search is model-gated — check the compatibility table for the exact model you are running before you depend on it.
The part that changes the whole calculation: prompt caching
Everything above assumes you pay full price for the overhead. With prompt caching you mostly do not — and this inverts a lot of the intuition about prompt length.
A cache read on Anthropic costs 0.1x the base input rate. A five-minute cache write costs 1.25x; a one-hour write costs 2x. Since a system prompt and a frozen tool list are the most perfectly stable prefix any application has, they are the ideal caching target — and the render order (tools, then system, then messages) means a single breakpoint on the last system block covers both.
response = client.messages.create(
model="claude-sonnet-5",
max_tokens=4096,
tools=TOOLS, # renders first, frozen, sorted
system=[{
"type": "text",
"text": SYSTEM_PROMPT,
# Covers tools + system in one entry.
"cache_control": {"type": "ephemeral"},
}],
messages=messages, # volatile — after the breakpoint
)
# Assert it worked. A silent cache miss looks exactly like success.
assert response.usage.cache_read_input_tokens > 0, "prefix is not stable"The arithmetic, as ratios
Take a 20,000-token fixed prefix served 1,000 times within the cache TTL, and express everything in units of "one uncached input token" so the numbers survive the next price change.
| Scenario | Cost per request (relative) | Cost over 1,000 requests | vs uncached |
|---|---|---|---|
| No caching | 1.0x × 20,000 | 20,000,000 units | — |
| 5m cache: 1 write + 999 reads | 1.25x once, then 0.1x | ~2,023,000 units | ~10x cheaper |
| 1h cache: 1 write + 999 reads | 2.0x once, then 0.1x | ~2,038,000 units | ~10x cheaper |
| Trimmed to 8,000 tokens, uncached | 1.0x × 8,000 | 8,000,000 units | ~2.5x cheaper |
Read the last two rows against each other. A 20,000-token cached prompt costs roughly a quarter of an 8,000-token uncached one. Aggressive trimming that costs you output quality is the wrong optimisation if caching is available to you — caching is worth more than trimming, by a wide margin, and it costs you nothing in quality because the model sees an identical prompt either way.
That reframes the whole exercise. The question is not "how short can I make this?" It is "how stable can I make this?" A large frozen prefix beats a small volatile one.
Where the caching story breaks
- Below the minimum, nothing caches and nothing tells you. The minimum cacheable prefix is model-specific and not monotonic: 1,024 tokens on Claude Opus 4.8 and Claude Sonnet 5, but 4,096 tokens on Claude Haiku 4.5. Trimming a prompt to 3,000 tokens makes it uncacheable on Haiku — a change that reads as an optimisation and behaves as a 10x regression.
- Traffic must be denser than the TTL. Requests arriving every eight minutes against a five-minute entry means every request is a 1.25x write and none is a read. You have made an uncached workload 25 percent more expensive.
- Anything volatile in the prefix kills it. A date, a user ID, a feature flag interpolated into the system prompt, or a tool list serialised from an unordered dict. Freeze the tool array and sort it deterministically.
- Concurrency cold-starts. An entry is only readable once the first response begins streaming, so N parallel requests against a cold prefix all pay the write. Send one, await first token, then fan out.
- OpenAI's contract differs. Caching is on by default with a 1,024-token minimum on GPT-5.6 and later (2,048 on older models), and a prefix stays eligible for 30 minutes after its last write or reuse. Cached tokens appear as usage.input_tokens_details.cached_tokens.
Deferred loading and caching together
These two techniques can look like they conflict, since one keeps definitions out of the prefix and the other depends on the prefix being stable. They do not. Deferred tools are excluded from the system-prompt prefix entirely, and when the model discovers one, the API appends a tool reference inline in the conversation rather than editing the prefix. The cached prefix is untouched. You get the small stable prefix and the large tool catalogue at the same time — which is precisely the combination that was impossible before.
Decision table
| Your situation | Do this first | Expect to trade |
|---|---|---|
| Low volume, small prompt | Nothing. Measure and move on. | Nothing — this is not your cost centre |
| High volume, stable prompt and tools | Turn on caching and assert on cache reads | Prompt rigidity: A/B tests now run cold |
| Prompt below the cache minimum | Consider padding with genuinely useful stable content | Slightly more context; verify quality is unharmed |
| Sparse, bursty traffic | 1h TTL, or accept no caching | 2x writes; needs two reads to break even |
| 50+ tools or multiple MCP servers | Deferred loading via tool search | A discovery round trip; description-quality now matters |
| One enormous auto-generated schema | Hand-write that one schema | Maintenance burden when the source changes |
| Per-user or per-tenant tool lists | Freeze one superset list; defer the rest | A slightly larger nominal list, far better cache economics |
What to do first
- Run the delta measurement above and write the four numbers down. Most teams are wrong about which of system prompt or tools is larger.
- Add a CI assertion that fixed overhead has not grown beyond a threshold. This is the single highest-leverage thing in the article, because bloat arrives gradually and nobody owns it.
- Turn on caching before you trim anything. Confirm cache_read_input_tokens is non-zero on the second identical request, and alert if the hit rate ever drops — a broken cache does not raise an error, it just quietly costs ten times more.
- Check your prefix against the cache minimum for every model you route to, including cheap fallbacks.
- Only now start trimming: kill session-derived parameters, collapse near-duplicate tools, shorten descriptions to the decision.
- If you are past roughly 30 tools or 10,000 tokens of definitions, move to deferred loading rather than trimming further — you are also fixing a selection-accuracy problem, not just a cost one.
Related reading on AI Workshack
- Prompt Caching Across Anthropic, OpenAI and Google: What Each One Actually Does — the full caching contract per provider, including the six cases where caching costs more than it saves.
- Context Window Management for Cost: Trimming, Summarising and Sliding Windows — the variable half of the same bill, and why trimming badly destroys the cache this article depends on.
- Why AI Agents Burn Tokens: The Agent Loop, Re-Sent Context, and Where the Spend Actually Goes — where fixed overhead is multiplied by loop iterations.
- Model Tiering and Routing: Matching Task Complexity to Model Cost — note that cache minimums differ per model, so routing and caching decisions are coupled.