The rule of thumb everyone quotes is wrong in the direction that costs you money. Here is what your provider is really metering, and how to measure it before the invoice arrives.
Every LLM bill is a token bill. You do not pay per request, per word, or per conversation — you pay per token, in and out. Which means the single most useful skill in cost control is not clever prompt engineering. It is being able to look at a request you are about to send and say, with reasonable confidence, how many tokens it contains.
Most teams cannot do this. They estimate with a rule of thumb, the estimate comes in 30 to 300 percent low, and the gap shows up as a bill nobody predicted. The gap is almost never the part of the prompt they wrote. It is the parts they never counted.
All prices in this article are as of August 2026 and are Anthropic first-party API rates. Model pricing changes, sometimes without a migration window. Verify current pricing against your provider's official pricing page before you commit a number to a budget, a contract, or a dashboard.A token is a vocabulary lookup, not a chunk of text
Language models do not read characters. They read integers. A tokenizer converts your text into a sequence of IDs, each one an entry in a fixed vocabulary built by compressing a large corpus into recurring subword fragments. Common English words map to a single token. Rare words split into pieces. Unusual byte sequences fall apart into fragments that carry almost no meaning.
This is why the tokenizer is a property of the model, not of your text. The same sentence produces different token counts on different models, because they were built against different vocabularies. There is no universal token.
Why "1 token ≈ 4 characters" misleads
Anthropic's own documentation offers the standard estimate: roughly 4 characters, or 0.75 words, per token in English. That number is honest and it is also an average over English prose. Your workload is probably not English prose. The rule degrades predictably, and it degrades in the expensive direction — real counts come in higher than the estimate, not lower.
| Content type | How the 4-chars-per-token rule behaves |
|---|---|
| English prose | Roughly accurate. This is what the average was measured on. |
| Source code | Underestimates. Indentation, punctuation, camelCase and snake_case identifiers, and operators fragment heavily. |
| JSON and structured data | Underestimates badly. Braces, quotes, colons, and repeated key names are each their own token, and keys repeat on every record. |
| Non-English text | Underestimates, often severely. Languages outside the vocabulary's training emphasis fragment toward per-character tokenization. |
| Logs, stack traces, UUIDs, hashes | Underestimates worst of all. High-entropy strings have no recurring fragments to compress against. |
| Numbers and tabular data | Underestimates. Long numeric strings split into small pieces. |
The practical consequence: if your workload is code, JSON, logs, or any non-English language, a character-based estimate is a floor, not a forecast. Treat it as the smallest the request could possibly be.
The token sources nobody counts
Here is the more expensive problem. Even a perfectly accurate character-to-token estimate of your prompt text will understate your bill, because a large fraction of what you are billed for is not prompt text at all. Every item below is metered as input tokens on every single request.
1. The system prompt
The system prompt is billed on every request, forever. It does not amortise. A 1,500-token system prompt on a service handling 200,000 requests a month is 300 million input tokens per month before a user has typed anything. People write system prompts once, treat them as configuration, and never look at the line again. It is usually the single largest fixed cost in the request.
2. Tool schemas, and the tool-use system prompt overhead
If you pass tools, you pay for the tool definitions — every name, description, and JSON schema property, on every request. That part is at least visible in your own code. The part that is not visible: the API automatically injects an additional system prompt that teaches the model how to use tools, and you are billed for it. It appears whenever at least one tool is present, and its size depends on the model and on the tool_choice setting you send. As of August 2026, the per-model overhead is:
| Model | tool_choice: auto or none | tool_choice: any or tool |
|---|---|---|
| Claude Opus 5 | 286 tokens | 406 tokens |
| Claude Opus 4.8 | 290 tokens | 410 tokens |
| Claude Opus 4.7 | 675 tokens | 804 tokens |
| Claude Sonnet 5 | 354 tokens | 474 tokens |
| Claude Sonnet 4.6 | 497 tokens | 589 tokens |
| Claude Haiku 4.5 | 496 tokens | 588 tokens |
Notice that this does not move in a straight line across versions. Opus 4.7 carries 675 tokens of overhead; Opus 5, a later and more capable model, carries 286. You cannot extrapolate this number from a model's tier or release order. It is a measured constant that changes with each model, which is exactly why you should re-measure it when you switch models rather than carrying forward an old figure.
Forcing a tool call also costs you. Moving tool_choice from auto to any or to a named tool adds roughly 90 to 130 tokens per request depending on the model. That is small in isolation and entirely real at volume — 120 extra tokens across a million requests is 120 million input tokens. The tradeoff is that forcing a tool is often the cheapest way to avoid a wasted turn where the model replies in prose instead of calling the tool you needed, and one avoided turn pays for a great many 120-token surcharges. Decide it on reliability, not on this line item, but know the line item exists.
3. Built-in tool definitions, which are much larger than you expect
Provider-defined tools ship their own definitions, and they are billed as input tokens on top of the tool-use system prompt above. As of August 2026:
| Tool | Additional input tokens |
|---|---|
| Bash tool (Opus 5, 4.8, 4.7) | 325 |
| Bash tool (Opus 4.6, Sonnet 4.6 and earlier) | 244 |
| Text editor (text_editor_20250429) | 700 |
| Computer use toolset (computer_toolset_20260801, default members) | ~4,500 |
| Browser use toolset (browser_toolset_20260801, default members) | ~6,600 |
Read the last row again. Declaring the browser toolset costs roughly 6,600 input tokens on every request in the loop — more than most teams' entire system prompt and tool schema budget combined. A computer-use or browser-use agent starts each turn several thousand tokens in debt before it has seen a single instruction. This is a real reason to disable toolset members you do not need; the configuration knobs exist precisely because the definitions are expensive.
4. Images, which are priced by area, not by file size
Image cost has nothing to do with the byte size of your PNG. Claude processes images in 28x28 pixel patches, and each patch is one visual token. The formula is exact:
visual_tokens = ceil(width / 28) * ceil(height / 28)A 1920x1080 screenshot is ceil(68.57) x ceil(38.57) = 69 x 39 = 2,691 visual tokens. Compressing that PNG harder changes the upload size and changes nothing about the bill.
There is a second trap here. Models have a maximum native resolution, and oversized images are downscaled before processing — which caps the token cost, but the cap differs by model generation. Claude 4.7 and later sit in a high-resolution tier (long edge up to 2576 px, up to 4,784 visual tokens); earlier models sit in a standard tier (1568 px, 1,568 visual tokens). The same 1920x1080 screenshot costs 2,691 visual tokens on a 4.7-or-later model and 1,560 on a standard-tier model, because the older tier downscales it to 1456x819 first.
So an agent that takes one screenshot per turn and runs 30 turns is spending roughly 80,700 visual tokens on images alone — about $0.40 at Opus 5 input rates, per run, before any text. The tradeoff is direct: downsampling images before sending cuts that cost proportionally and costs you fidelity on small text and dense UI. If your agent is reading dense documents or clicking small controls, pay for the resolution. If it is checking whether a page loaded, do not.
5. PDFs and documents
A PDF sent as a document block is converted into text and images before the model sees it — page images included. A long PDF is therefore both a large text cost and a large visual-token cost simultaneously. Page count is a far better cost predictor than file size, and a scanned PDF costs dramatically more than a text-native one of the same length because every page arrives as an image.
6. Web-fetched pages and search results
Anything a server-side tool pulls into the conversation becomes input tokens, and it stays in the conversation for every subsequent turn. Anthropic publishes these rules of thumb:
| Fetched content | Approximate tokens | Cost at Opus 5 input ($5/MTok) |
|---|---|---|
| Average web page (10 kB) | ~2,500 | $0.0125 |
| Large documentation page (100 kB) | ~25,000 | $0.125 |
| Research paper PDF (500 kB) | ~125,000 | $0.625 |
Web fetch carries no surcharge beyond those tokens. Web search does: $10 per 1,000 searches, on top of the tokens the results generate. That is a fixed per-call fee independent of model, which means it is a rounding error on a Fable 5 request and a meaningful share of a Haiku 4.5 request. A research agent that fetches six documentation pages has quietly added 150,000 input tokens to its context, and it will re-send them on every turn thereafter.
Cap this at the source. The web fetch tool accepts a max_content_tokens parameter. Set it. Without a cap, a single unexpectedly large page can add six figures of tokens to a request you thought was small.7. Thinking and reasoning output
Extended thinking is billed as output tokens — the expensive category — and it is billed whether or not you display it. On models where thinking is on by default, a request you believe produces a 200-token answer may bill for several thousand output tokens. Turning off the display does not turn off the billing; it only hides the evidence. The effort setting is the actual cost lever here.
8. The conversation history you re-send
The API is stateless. Every turn re-sends the entire prior conversation as input tokens. This is the dominant cost in any multi-turn or agentic workload and it deserves more than a paragraph — see the separate article on agent token burn in this track for the full mechanism.
Counting properly: the count_tokens endpoint
Every provider that charges per token gives you a way to count them exactly. For Claude it is the /v1/messages/count_tokens endpoint, exposed in the SDKs as client.messages.count_tokens(...). It runs the model's real tokenizer against your real request and returns the number you will be billed on. It is not an estimate.
The critical detail is that it accepts the whole request shape — system prompt, tools, and messages — so it captures the hidden overhead described above, not just your prompt text:
from anthropic import Anthropic
client = Anthropic()
result = client.messages.count_tokens(
model="claude-opus-5", # counts are model-specific
system=SYSTEM_PROMPT,
tools=TOOL_DEFINITIONS, # includes the tool-use overhead
messages=messages,
)
print(result.input_tokens)
Count the request you are actually going to send. Counting only the user message tells you the least interesting number in the request.
Why tiktoken is the wrong tool for Claude
This mistake is everywhere and it is worth being blunt about. tiktoken is OpenAI's tokenizer. It implements OpenAI's vocabularies. It knows nothing about Claude, and it will happily return a confident number for Claude input that is simply not the number you will be charged for. It appears in Claude codebases because it is the tokenizer library most people met first, it installs cleanly, and it never errors.
Because tokenizers are model-specific, a count from the wrong tokenizer is not a rough estimate — it is a different measurement. Anthropic's guidance is explicit: tiktoken undercounts Claude tokens by roughly 15 to 20 percent on typical text, and by considerably more on code and non-English input. If you have built cost dashboards, context-window guards, or chunking logic on tiktoken counts for a Claude workload, every one of those numbers is low, and the guards will fail on exactly the inputs — dense code, foreign-language documents — where failing is most expensive.
The same principle runs in both directions. Do not use Claude's endpoint to size an OpenAI or Google request either. Count with the tokenizer of the model you are actually billing against.
The tokenizer-generation trap
This is the subtlest item on the list, and in August 2026 it is catching a lot of teams out.
Claude 4.7 and later models use a newer tokenizer that produces approximately 30 percent more tokens for the same text than Sonnet 4.6 and earlier. The exact increase depends on content and workload shape. Token counts are not comparable across tokenizer generations.Sit with what that means. The tokenizer changed as part of a capability improvement — finer-grained tokens help the model — but it silently redefines the unit your entire cost model is denominated in. Identical text, identical prompt, newer model: about 30 percent more tokens.
The failure modes that follow are all quiet ones:
- Cost-per-request dashboards show a jump after a model upgrade and get blamed on prompt bloat that never happened.
- Token-based alert thresholds tuned on an older model fire constantly, and get raised until they no longer catch anything.
- Chunking logic sized to fit a context budget starts producing chunks that overflow.
- A model comparison run on token counts ranks the newer model as more expensive per unit of work when it may well be cheaper.
- Context-window headroom you thought you had disappears — the window is the same size, but your text now occupies more of it.
The correction is not complicated, but it has to be deliberate:
- Re-baseline on every tokenizer-generation change. Run your real prompts through count_tokens on the new model and record fresh constants. Do not scale the old numbers by a rule of thumb.
- Stamp token metrics with the model that produced them. A token count without a model ID is not a measurement, and comparing two of them is meaningless.
- Compare models in dollars per completed task, never in tokens. Dollars are commensurable across tokenizer generations; tokens are not.
- Re-tune thresholds, chunk sizes, and context guards as part of the migration, not as a follow-up ticket after they start misfiring.
If you keep one sentence from this article, keep this one: a token is not a fixed unit of text. It is a unit defined by a specific model's tokenizer, and it can be redefined by a model upgrade you did not think of as a pricing change.Building a pre-flight cost estimate
Put the pieces together and you get something genuinely useful: a cost estimate before you send the request. The structure is always the same — a fixed floor that every request pays, plus a variable part, plus an expected output.
Step 1: measure the fixed floor once
Your floor is the system prompt, the tool schemas, the tool-use system prompt overhead, and any built-in tool definitions. It does not change between requests, so measure it once per prompt version and store the number. A worked example for a modest agent on Opus 5:
| Component | Tokens |
|---|---|
| System prompt | 1,200 |
| Six tool schemas | 2,400 |
| Tool-use system prompt (Opus 5, tool_choice auto) | 286 |
| Bash tool definition | 325 |
| Fixed floor per request | 4,211 |
At $5 per million input tokens that is about $0.021 per request in fixed overhead. Across 100,000 requests a month, roughly $2,100 — paid before any user input is processed. Served from a warm prompt cache at the 0.1x read rate, the same floor costs roughly $210. That single comparison is the entire argument for prompt caching, and it is covered properly in this track's caching article.
Step 2: estimate the variable input and the expected output
Variable input is the user's message plus any retrieved documents, images, or fetched pages. Expected output is the part people skip, and it is the part that dominates the bill — output is priced at 5x input on every current Claude tier. Estimate it from measured history: run 50 real requests, record the actual output_tokens, and use the median and the 95th percentile. Never use your intuition about how long the answer "should" be.
Step 3: write it down as code
from anthropic import Anthropic
client = Anthropic()
# Rates as of August 2026, USD per million tokens. Verify before trusting.
PRICES = {
"claude-opus-5": {"in": 5.00, "out": 25.00},
"claude-sonnet-5": {"in": 2.00, "out": 10.00},
"claude-haiku-4-5": {"in": 1.00, "out": 5.00},
}
def preflight_estimate(model, system, tools, messages, expected_output_tokens):
counted = client.messages.count_tokens(
model=model, system=system, tools=tools, messages=messages,
).input_tokens
rate = PRICES[model]
input_cost = counted * rate["in"] / 1_000_000
output_cost = expected_output_tokens * rate["out"] / 1_000_000
return {
"input_tokens": counted,
"expected_output_tokens": expected_output_tokens,
"input_cost": input_cost,
"output_cost": output_cost,
"total": input_cost + output_cost,
}
Two guardrails on using this. First, the estimate excludes server-tool surcharges such as web search, so add those separately if your request can trigger them. Second, it says nothing about caching — it counts tokens, not which price band they will land in. For that you need the actual usage object after the fact, which is the subject of the companion article in this track.
The tradeoff: counting is not free
The counting endpoint is a network round trip. It has its own latency and its own rate limit. Calling it on every request in a latency-sensitive path buys you an accurate number at the cost of a slower product, and it is rarely worth it.
The pattern that works: call it in development and in CI to establish per-prompt-shape constants, and at runtime only where the decision actually depends on the answer — gating an oversized user upload, choosing between a cheap and an expensive model, or refusing a request that would blow a per-user budget. Everywhere else, estimate from the stored constants and reconcile against the real usage object after the response comes back. That is free, exact, and available on every request.
What to actually do this week
- Count your system prompt and tool schemas with count_tokens. Multiply by your monthly request volume. Most teams find a number they did not expect.
- Grep your codebase for tiktoken. If it is being used against a Claude workload, every count downstream of it is low by at least 15 to 20 percent.
- Check which tokenizer generation your token metrics were baselined on, and stamp the model ID onto every token metric you store.
- Add max_content_tokens to any web fetch tool you have deployed without one.
- Record actual output_tokens for 50 real requests and replace your intuition about response length with a measured median and p95.
None of this is optimisation. It is measurement — the step that has to come first, because you cannot reduce a number you have never accurately counted. Once you have the counts, the rest of this track covers what to do with them: reading the usage object and turning it into per-request cost, prompt caching, batch processing, model routing, output-token control, agent token burn, and monitoring.
All prices in this article are as of August 2026 and are Anthropic first-party API rates. Model pricing changes, sometimes without a migration window. Verify current pricing against your provider's official pricing page before you commit a number to a budget, a contract, or a dashboard.