Four counters, four different prices, and a set of multipliers that stack. Here is how to turn a usage object into a cost number you can actually track per request.
Your provider's dashboard gives you a monthly total. That number is useless for making decisions. It cannot tell you which feature is expensive, whether your caching is working, or whether last Tuesday's spike was more traffic or worse behaviour.
The number you need is cost per request, and every API response already contains everything required to compute it. The obstacle is that most people read the usage object as if it had two fields. It has at least four, they are priced very differently from each other, and they do not sum the way you would assume.
All prices in this article are as of August 2026 and are Anthropic first-party API rates. Model pricing changes, sometimes without a migration window. Verify current pricing against your provider's official pricing page before you commit a number to a budget, a contract, or a dashboard.First principle: output tokens cost 5x input
Before decomposing anything, internalise the ratio that governs every decision downstream. As of August 2026, across the entire current Claude lineup:
| Model | Input $/MTok | Output $/MTok | Output-to-input ratio |
|---|---|---|---|
| Claude Fable 5 | $10 | $50 | 5x |
| Claude Opus 5 | $5 | $25 | 5x |
| Claude Sonnet 5 | $2 | $10 | 5x |
| Claude Haiku 4.5 | $1 | $5 | 5x |
The ratio is 5x on every tier. That is not a coincidence you can ignore — it is a structural fact you can reason from, and it survives moving between models. One output token costs the same as five input tokens. Which gives you a break-even you can hold in your head:
Output tokens dominate your bill whenever output_tokens exceeds input_tokens / 5. Once prompt caching is working and your input is arriving at the 0.1x cache-read rate, that threshold drops to input_tokens / 50 — a cached input token is fifty times cheaper than an output token.Two concrete shapes, both on Opus 5, show how sharply this cuts:
| Request shape | Input cost | Output cost | Where the money is |
|---|---|---|---|
| Summarise a document: 20,000 in, 500 out | $0.100 | $0.0125 | Input, 89% of cost |
| Agent turn, cached history: 20,000 cache-read, 2,000 out | $0.010 | $0.050 | Output, 83% of cost |
Same model, same order-of-magnitude token volume, and the levers that matter are completely different. In the first case you optimise the input — trim the document, cache the prefix. In the second, trimming input is close to pointless and the only lever that matters is response length. Guessing which case you are in is how teams spend a quarter optimising the wrong half of the bill. This track has a dedicated article on controlling output tokens; the point here is knowing when to reach for it.
The four-way split in the usage object
Here is a real usage object shape, the kind returned by a cached request that also ran a web search:
{
"usage": {
"input_tokens": 105,
"output_tokens": 6039,
"cache_read_input_tokens": 7123,
"cache_creation_input_tokens": 7345,
"server_tool_use": {
"web_search_requests": 1
}
}
}
Four token counters, four price bands. What each one means:
| Field | What it counts | Price, relative to base input |
|---|---|---|
| input_tokens | Input tokens processed fresh this request — the uncached remainder only | 1x (base rate) |
| cache_creation_input_tokens | Input tokens written into the prompt cache this request | 1.25x for a 5-minute TTL; 2x for a 1-hour TTL |
| cache_read_input_tokens | Input tokens served from an existing cache entry | 0.1x |
| output_tokens | Everything the model generated, including thinking tokens | 5x base input, per the model's output rate |
The trap: input_tokens is not your total input. The three input counters are separate and additive. Total input tokens = input_tokens + cache_creation_input_tokens + cache_read_input_tokens. Code that computes cost as input_tokens * input_rate + output_tokens * output_rate silently omits every cached token from the bill — and the better your caching works, the larger the omission. Teams have watched their computed costs drop after enabling caching by far more than caching could possibly save, because the cached tokens fell out of the calculation entirely.A useful sanity check follows from that: if you enable caching and your computed cost per request drops by more than about 90 percent, you have a summing bug, not a win. The floor is 0.1x on the cached portion, and only on the cached portion.
The cache break-even, in one line
Writing to cache costs more than not caching. A 5-minute write is 1.25x base input; a read is 0.1x. So a 5-minute cache entry pays for itself after a single read, and a 1-hour entry (2x write) after two. Anything beyond that is profit. This is why cache hit rate is the metric to watch rather than raw cache token volume — a cache you write and never read is a 25 percent surcharge. The caching article in this track covers placement and invalidation properly.
A worked cost calculation, from that usage object
Take the usage object above and price it on Opus 5, assuming a 5-minute cache TTL (write rate $6.25/MTok, read rate $0.50/MTok):
| Line item | Calculation | Cost |
|---|---|---|
| Fresh input | 105 x $5 / 1M | $0.000525 |
| Cache writes | 7,345 x $6.25 / 1M | $0.045906 |
| Cache reads | 7,123 x $0.50 / 1M | $0.003562 |
| Output | 6,039 x $25 / 1M | $0.150975 |
| Web search | 1 search x $10 / 1,000 | $0.010000 |
| Total | $0.210968 |
Now read what that table is telling you. The request moved 20,612 tokens in total, and output was 6,039 of them — 29 percent of the tokens. But output accounts for $0.151 of $0.211, which is 72 percent of the cost. Token volume and cost are not the same distribution, and a dashboard that charts token counts will point you at the wrong problem.
Second observation: the cache write cost $0.046 on this request, roughly thirteen times what the cache read saved. That is normal and correct on the request that populates the cache. It only becomes a problem if the reads never arrive.
The same request on Sonnet 5
Repricing the identical usage object at Sonnet 5 rates ($2/$10, write $2.50, read $0.20) gives a total of about $0.0904 — roughly 2.3x cheaper. But look at what does not scale:
| Component | Opus 5 | Share of total | Sonnet 5 | Share of total |
|---|---|---|---|---|
| Token costs | $0.20097 | 95.3% | $0.08039 | 88.9% |
| Web search fee | $0.01000 | 4.7% | $0.01000 | 11.1% |
| Total | $0.21097 | $0.09039 |
The web search fee is model-independent. It is 4.7 percent of the Opus 5 request and 11.1 percent of the Sonnet 5 one. This is the general shape of the thing: as you route work down to cheaper models, fixed per-call fees become a progressively larger share of what you spend, and eventually they set a floor that model choice cannot get you under. If your cheap-tier requests each fire three searches, you are paying $0.03 in search fees on a request whose tokens cost less than that. Model routing has a floor, and this is where it lives — see the model routing article in this track for the full decision framework.
Note also that this usage object comes from a real tokenizer generation. Repricing it against a model on a different tokenizer generation is not a like-for-like comparison: the same source text would produce a different token count. Compare dollars per completed task, not dollars per token.
How the multipliers stack
On top of the four base categories sit several multipliers, and they compose. The order of operations, as of August 2026:
- Start with the model's base input and output rates.
- Apply the cache multiplier to whichever tokens fall in that band: 1.25x for a 5-minute write, 2x for a 1-hour write, 0.1x for a read.
- Apply the Batch API discount if the request went through batch: 50% off input and output, and off the cache bands too.
- Apply the data-residency multiplier if inference_geo is set to "us": 1.1x on all token categories, including cache writes and cache reads.
Worked out for Opus 5, in effective dollars per million tokens:
| Category | Standard | Batch (50% off) | inference_geo us (1.1x) | Batch + us |
|---|---|---|---|---|
| Input | $5.00 | $2.50 | $5.50 | $2.75 |
| Output | $25.00 | $12.50 | $27.50 | $13.75 |
| 5-min cache write | $6.25 | $3.13 | $6.88 | $3.44 |
| 1-hour cache write | $10.00 | $5.00 | $11.00 | $5.50 |
| Cache read | $0.50 | $0.25 | $0.55 | $0.28 |
Three things to take from that grid:
- Batch and caching combine. They are not alternatives, and using both is materially cheaper than either alone.
- The 1M-token context window on 4.6 and later models bills at standard rates — there is no long-context surcharge. A 900k-token request costs the same per token as a 9k one. Context length is a cost problem only because of volume, not because of a rate penalty.
- inference_geo: "us" applies to every category, cache included. It is a flat 10 percent on your whole token bill, which is worth knowing before someone enables it for a compliance reason nobody costed.
Multipliers have exclusions, and they are the kind that surprise you at invoice time. The Batch API discount does not apply to Managed Agents sessions — sessions are stateful and interactive, so there is no batch mode. Fast mode is not available with batch at all. And Managed Agents bills session runtime as a separate line item on top of tokens. Check exclusions before you build a forecast on a discount.The batch tradeoff is the obvious one and worth stating: 50 percent off in exchange for asynchronous processing with a completion window measured in hours. Anything a user is waiting on is disqualified. The batch article in this track covers where the line sits.
Turning usage into cost-per-request you can track
The mechanics are straightforward once the four-way split is clear. What matters is what you store alongside the number.
# Rates as of August 2026, USD per million tokens. Verify before trusting.
PRICE_TABLE_VERSION = "2026-08-25"
PRICES = {
"claude-fable-5": {"in": 10.00, "out": 50.00},
"claude-opus-5": {"in": 5.00, "out": 25.00},
"claude-sonnet-5": {"in": 2.00, "out": 10.00},
"claude-haiku-4-5": {"in": 1.00, "out": 5.00},
}
CACHE_WRITE_MULTIPLIER = {"5m": 1.25, "1h": 2.00}
CACHE_READ_MULTIPLIER = 0.10
WEB_SEARCH_PER_CALL = 10.00 / 1000
def request_cost(model, usage, cache_ttl="5m", batch=False, us_only=False):
rate = PRICES[model]
base_in, base_out = rate["in"], rate["out"]
fresh = usage.input_tokens * base_in
written = (usage.cache_creation_input_tokens or 0) * base_in \
* CACHE_WRITE_MULTIPLIER[cache_ttl]
read = (usage.cache_read_input_tokens or 0) * base_in * CACHE_READ_MULTIPLIER
generated = usage.output_tokens * base_out
token_cost = (fresh + written + read + generated) / 1_000_000
if batch:
token_cost *= 0.50
if us_only:
token_cost *= 1.10
searches = getattr(usage, "server_tool_use", None)
search_cost = 0.0
if searches is not None:
search_cost = (getattr(searches, "web_search_requests", 0) or 0) \
* WEB_SEARCH_PER_CALL
return token_cost + search_cost
What to log on every request
The cost number alone ages badly. Log the inputs to it as well, so that when rates change or you find a bug in the calculation you can reprice your entire history instead of throwing it away:
- The model ID — without it, token counts are not comparable, because tokenizer generations differ.
- All four token counters, stored raw and separately. Never store a pre-summed input total.
- Any server_tool_use counters, especially web_search_requests, since those are priced outside the token model.
- The feature flags that changed the price: batch, inference_geo, cache TTL, fast mode.
- The computed cost, and the version of the price table used to compute it.
- A dimension you can group by — user, tenant, feature, endpoint. Cost per request is only actionable when you can slice it. The cost attribution article in this track goes further on this.
The tradeoff: your number will not match the invoice
A client-side cost calculation is a model of your bill, not your bill. It will drift, for legitimate reasons: negotiated or volume discounts you did not encode, free-tier allowances, provider-side rounding, failed requests that were or were not billed, and rate changes that land mid-month. Requests that error partway through can bill for partial output that your instrumentation never sees.
Do not try to close that gap. Reconcile monthly against the provider's actual invoice, record the ratio between your computed total and the real one, and if that ratio is stable you have a perfectly good relative signal — which is what you need for comparing features, users, and prompt versions. If the ratio starts moving, that is your bug, and it is worth chasing. Treat your computed cost as a high-resolution leading indicator, and the invoice as the ground truth.
The three numbers worth watching
Once per-request cost is flowing, most of the value comes from three derived metrics rather than the raw total:
- Output share of cost. Rising output share means responses are getting longer, thinking is running deeper, or your caching started working — all of which change which lever to reach for next.
- Cache hit rate, measured as cache_read_input_tokens divided by total input tokens. If this is near zero while cache_creation is large, you are paying the 1.25x write surcharge for nothing, and something is silently invalidating your prefix.
- Cost per completed task, not cost per request. An agent that needs eight cheap requests to finish a job is not cheaper than one that needs two expensive ones. This is the only metric that survives a model change, a tokenizer change, and a prompt rewrite.
The monitoring article in this track covers instrumentation and alerting on these in depth.
If you are on a subscription plan, not the API
You will not see a usage object, and you are not billed per token. But the same arithmetic explains what consumes your limits, because the underlying resource is identical.
- Long responses cost roughly five times what the equivalent length of input costs. Asking for a summary is far cheaper than asking for a rewritten document.
- Deep reasoning is output. A request that thinks hard consumes far more of your allowance than its visible answer suggests.
- Every turn in a long conversation re-sends the whole conversation. Starting a fresh conversation for an unrelated question is the single highest-leverage habit available to you.
- Attaching a large PDF or a full-resolution screenshot spends allowance on every subsequent turn in that conversation, not just the one where you attached it.
The instrumentation differs. The economics do not.
The short version
- Output is 5x input on every current tier. Find out which side of that ratio your workload sits on before optimising anything.
- There are four token counters, not two, and the three input counters are separate and additive. Summing them wrongly is the most common cost-instrumentation bug there is.
- Multipliers stack in a defined order: base rate, then cache band, then batch, then data residency. Check exclusions before forecasting a discount.
- Compute cost per request from the raw usage object, store the counters alongside the result, and reconcile against the invoice monthly.
- Track output share, cache hit rate, and cost per completed task. Those three tell you what a monthly total never will.
All prices in this article are as of August 2026 and are Anthropic first-party API rates. Model pricing changes, sometimes without a migration window. Verify current pricing against your provider's official pricing page before you commit a number to a budget, a contract, or a dashboard.