Three providers, three completely different caching contracts — and the cases where turning caching on makes your bill go up.

If you send the same 40,000-token system prompt on every request, you are paying for those 40,000 tokens every single time. Prompt caching is the mechanism that stops that. All three major providers offer it, all three advertise roughly the same headline discount, and all three implement it so differently that code and intuition do not transfer between them.

This article covers the mechanism, not the marketing: who decides what gets cached, what the write costs, how long the entry survives, what silently breaks it, and — the part nobody writes about — the specific situations where enabling caching makes your bill larger.

All prices, multipliers and minimums in this article were verified against provider documentation as of August 2026. Provider pricing changes frequently — verify current pricing and limits against the provider's own docs before you build a business case on them.

The one thing all three have in common: prefix matching

Every implementation caches a prefix, not a set of chunks. The provider hashes your request from the beginning up to some point, and if that hash matches an existing entry, everything up to that point is served from cache. Nothing else about the mechanism is shared between providers, but this part is universal and it drives every design decision you will make.

Two consequences follow immediately. First, one changed byte anywhere in the prefix invalidates everything after it — a timestamp at the top of your system prompt destroys the cache for the entire request. Second, the order of your prompt is now load-bearing. Static content has to come first; volatile content has to come last. If your prompt is currently organised by topic, caching will force you to reorganise it by volatility.

Anthropic: explicit breakpoints you place yourself

Anthropic gives you the most direct control and the most responsibility. You mark specific content blocks with a `cache_control` field, and each mark creates a cache entry covering the entire prefix up to and including that block. You get up to four breakpoints per request.

import anthropic
 
client = anthropic.Anthropic()
 
response = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    system=[
        {
            "type": "text",
            "text": STABLE_INSTRUCTIONS + LARGE_REFERENCE_DOCUMENT,
            # Everything from the start of the request up to and including
            # this block becomes one cache entry.
            "cache_control": {"type": "ephemeral", "ttl": "5m"},
        }
    ],
    messages=[{"role": "user", "content": user_question}],  # volatile — after the breakpoint
)
 
print(response.usage.cache_creation_input_tokens)  # tokens written this request
print(response.usage.cache_read_input_tokens)      # tokens served from cache
print(response.usage.input_tokens)                 # uncached tokens, full price

If you do not want to think about placement, there is a top-level form — passing `cache_control={"type": "ephemeral"}` as a request-level parameter tells the API to manage breakpoints for you as the conversation grows. Start there; move to explicit block-level breakpoints only when you have a specific prefix you need to pin.

The prefix order is fixed

Anthropic assembles the cacheable prefix in a fixed order: `tools`, then `system`, then `messages`. Each level invalidates the ones after it. Change a tool name, description, or schema and you have invalidated the tool definitions, the system prompt, and the entire message history in one go. This is the single most common cause of a cache hit rate that mysteriously sits at zero in an agentic app — a tool list assembled from an unordered dictionary produces a different byte sequence on every process restart.

Change What it invalidates
Tool definitions (names, descriptions, schemas) tools + system + messages — everything
`tool_choice` parameter messages only
Adding or removing an image anywhere messages only
Toggling web search or citations system + messages
Providing tool results as user messages Nothing — cache stays valid

Write cost, read cost, and the break-even point

This is the number that decides whether caching is worth it, and it is a multiplier on the base input price rather than an absolute figure — which makes it far more durable than any dollar amount.

Operation Multiplier on base input price Entry lifetime
5-minute cache write 1.25x 5 minutes
1-hour cache write 2x 1 hour
Cache read (hit) 0.1x Same as the preceding write

Writing costs more than not caching. That is the part people skip. A 5-minute write costs 1.25x, so it pays for itself after a single read. A 1-hour write costs 2x, so it needs two reads before you are ahead. Anthropic's own documentation states this break-even explicitly. Reading the cache refreshes its TTL at no extra charge, so a steadily-used entry can live far longer than its nominal window without being rewritten.

Minimum cacheable size varies by model — and failing it is silent

A prefix below the minimum will not cache, and you get no error. The request succeeds, `cache_read_input_tokens` stays at zero, and you keep paying full price forever.

Model Minimum cacheable prefix
Claude Opus 4.8 (`claude-opus-5`) 1,024 tokens
Claude Sonnet 5 (`claude-sonnet-5`) 1,024 tokens
Claude Haiku 4.5 (`claude-haiku-4-5`) 4,096 tokens
The Haiku threshold catches teams out constantly. A prompt that caches perfectly on Sonnet 5 at 2,000 tokens will silently refuse to cache on Haiku 4.5, which needs 4,096. If you route the same prompt across model tiers, check the minimum for every tier you route to.

OpenAI: automatic by default, with an explicit override

OpenAI inverts the model. Caching is on by default and you do not place breakpoints — the platform places one at the end of the latest eligible message. There is an explicit mode (`prompt_cache_breakpoint` with `mode` set to `explicit`) if you need to control placement yourself, but the default path requires no code change at all.

The eligibility threshold is 1,024 tokens on GPT-5.6 and later, and 2,048 tokens on older models. What gets cached is the model's full rendered context — OpenAI's own injected instructions, your developer messages, tool definitions, and conversation history including text, images and documents.

from openai import OpenAI
 
client = OpenAI()
 
response = client.responses.create(
    model="gpt-5.6-sol",
    input=[
        {"role": "developer", "content": STABLE_INSTRUCTIONS + LARGE_REFERENCE_DOCUMENT},
        {"role": "user", "content": user_question},
    ],
    # Routes requests that share a prefix to the same cache. Use a stable key
    # per tenant / per prompt-version, NOT per request.
    prompt_cache_key="support-triage-v7",
)
 
usage = response.usage
print(usage.input_tokens_details.cached_tokens)      # tokens served from cache
print(usage.input_tokens_details.cache_write_tokens)  # tokens written this request

Pricing on GPT-5.6 and later mirrors Anthropic's shape closely: cache writes are billed at 1.25x the standard uncached input rate, and subsequent reads at 0.1x. Retention differs, though. On GPT-5.6 and later, `prompt_cache_options.ttl` accepts `"30m"` — which is both the default and, at time of writing, the only supported value — and the prefix stays eligible for 30 minutes after its most recent write or reuse. Earlier models use a different parameter, `prompt_cache_retention`, with values `"in_memory"` or `"24h"`.

One implementation detail worth knowing: OpenAI's cache stores key-value tensors, not the tokens themselves. This is why the cache is model-specific and why a model version bump wipes it — the stored artefact is tied to the weights.

`prompt_cache_key` is a routing hint, not a cache identifier. Give it a value that is stable across many requests — a prompt version, a tenant ID, a workload name. A per-request UUID here is worse than passing nothing, because it scatters requests that share a prefix across different cache shards.

Google: implicit for free, explicit for control (and a storage bill)

Google runs both modes side by side, and they have genuinely different economics. Implicit caching is on by default for Gemini 2.5 and later, requires no code, and follows the same prefix logic as everyone else. Minimums are 4,096 tokens on Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash, and 2,048 tokens on Gemini 2.5 Flash and 2.5 Pro. Google's guidance for maximising hits is exactly what you would expect: put large, common content at the start of the prompt, and send similar-prefix requests close together in time.

Explicit caching is the outlier across all three providers. You create a named cache object as a first-class resource, then reference it by name on subsequent generation calls.

from google import genai
from google.genai import types
 
client = genai.Client()
 
# Create a cache object. If ttl is omitted it defaults to 1 hour.
cache = client.caches.create(
    model="gemini-3.6-flash",
    config=types.CreateCachedContentConfig(
        display_name="policy-handbook-v3",
        system_instruction=STABLE_INSTRUCTIONS,
        contents=[LARGE_REFERENCE_DOCUMENT],
        ttl="3600s",
    ),
)
 
response = client.models.generate_content(
    model="gemini-3.6-flash",
    contents=user_question,
    config=types.GenerateContentConfig(cached_content=cache.name),
)
 
# Extend, or clean up when the work is done — storage bills by wall-clock time.
client.caches.update(name=cache.name, config=types.UpdateCachedContentConfig(ttl="600s"))
client.caches.delete(cache.name)

Cached input tokens are billed at roughly one tenth of the standard input rate, matching the other two providers. The difference is that Google also charges storage, per million tokens per hour, for as long as the cache object exists. On Gemini 3.7 and 3.6 Flash that storage rate is $0.50 per million tokens per hour (rising to $1.00 from 1 January 2027 per Google's published schedule); on Gemini 3.5 Flash it is $1.00; on Gemini 2.5 Pro it is $4.50.

Google's explicit cache storage accrues on wall-clock time whether or not you ever read it. A 200,000-token cache held on Gemini 2.5 Pro for a day costs roughly 0.2 x $4.50 x 24 — about $21.60 — before a single request touches it. Always set a TTL that matches your actual traffic pattern, and delete caches when the job ends. Verify these rates before relying on them; they are dated August 2026.

Note also that Google's Interactions API supports implicit caching only. Explicit cache objects require the generateContent API. And per Google's own docs, "the model doesn't make any distinction between cached tokens and regular input tokens" — caching is purely a billing and latency mechanism, not a quality one, on all three platforms.

Side by side

Anthropic OpenAI Google
Default state Off — you opt in On automatically Implicit on; explicit opt-in
Who places the breakpoint You (up to 4), or top-level auto Platform, or explicit mode Platform (implicit); you (explicit)
Minimum prefix 1,024–4,096 tokens by model 1,024 (GPT-5.6+) / 2,048 (older) 2,048–4,096 tokens by model
Write cost 1.25x (5m) / 2x (1h) 1.25x (GPT-5.6+) No separate write charge
Read cost 0.1x base input 0.1x base input ~0.1x base input
Lifetime 5m or 1h, refreshed free on use 30m from last write or reuse 1h default, you set the TTL
Storage charge None None Yes, per MTok per hour (explicit)
Cache lifecycle object None — implicit entries None — implicit entries Named, listable, updatable, deletable

When caching costs you more than it saves

Every write is more expensive than not caching at all. Caching is a bet that the write will be amortised by reads. Here are the six situations where that bet loses.

  1. **Your prefix is below the minimum.** No error, no warning, no cache. You keep paying full price and you think the feature is on. This is the most common failure and the easiest to check: if `cache_read_input_tokens` (or `cached_tokens`) is zero across repeated identical requests, caching is not happening.
  2. **Write-once, never read.** A one-off document analysis with a single request pays the 1.25x write multiplier and gets nothing back. On a 1-hour Anthropic TTL that is a flat 2x on your input tokens for zero benefit.
  3. **Your request interval exceeds the TTL.** A five-minute Anthropic entry with traffic arriving every eight minutes means every single request is a write and none is a read. You have converted a 1.0x workload into a 1.25x workload. Either move to the 1-hour TTL — accepting the 2x write and needing two reads to break even — or accept that this workload is not cacheable.
  4. **Volatile content sits inside the prefix.** A current timestamp, a request ID, a randomised few-shot sample, an unsorted JSON serialisation. Each request writes a new entry and reads nothing. You pay the write multiplier permanently.
  5. **Your tool list is not deterministic.** On Anthropic, tool definitions sit at the front of the prefix, so any instability there invalidates everything downstream. Tools assembled from a set, a dict iteration, or a plugin registry that loads in filesystem order will produce different bytes across processes. Sort them, freeze them, and treat the serialised tool array as a versioned artefact.
  6. **Cold-start stampede.** Cache entries only become readable once the first response begins. Firing fifty parallel requests at a cold prefix means fifty concurrent writes at 1.25x and zero reads. Send one request first, wait for it, then fan out.
A quick sanity rule: caching pays when the same prefix is read more times than the write multiplier costs, within the TTL. For a 5-minute Anthropic entry that is one read. For a 1-hour entry it is two. For Google explicit caching you must also clear the storage charge, which scales with cache size and holding time rather than with request count.

The tradeoffs nobody puts in the pricing page

Caching does not cost you output quality — the model sees an identical prompt either way. It costs you three other things, and they are real.

  • **Prompt rigidity.** Once a prefix is cached, changing it is expensive. A/B testing two system prompts means maintaining two warm caches, or accepting that the experiment arm runs cold. Teams that cache aggressively iterate on prompts more slowly, and that is a genuine product cost.
  • **Ordering discipline you now have to enforce.** Static-before-volatile is not a natural way to write a prompt. It has to be documented, and it has to survive the next person who adds a personalisation line at the top of the system prompt because that felt like the right place for it.
  • **A new class of silent failure.** A broken cache does not throw. It shows up weeks later as a cost line that never came down. The only defence is to assert on it: log the cache-read field on every request, alert when the hit rate drops below a threshold, and treat a zero hit rate as an incident rather than a curiosity.

Latency is the one place caching is unambiguously positive — a cache read skips prefill for that portion of the prompt, so time-to-first-token improves alongside cost. If your workload is latency-sensitive, that may matter more than the money.

How to verify it is actually working

Do not trust that caching is on because you set the parameter. Assert it. Every provider returns the numbers you need in the usage object.

import anthropic
 
client = anthropic.Anthropic()
 
def cached_call(user_question: str):
    response = client.messages.create(
        model="claude-sonnet-5",
        max_tokens=1024,
        system=[{
            "type": "text",
            "text": STABLE_INSTRUCTIONS + LARGE_REFERENCE_DOCUMENT,
            "cache_control": {"type": "ephemeral"},
        }],
        messages=[{"role": "user", "content": user_question}],
    )
    u = response.usage
    total_input = u.cache_read_input_tokens + u.cache_creation_input_tokens + u.input_tokens
    hit_rate = u.cache_read_input_tokens / total_input if total_input else 0.0
    # Emit this as a metric, not a print. A hit rate that drifts to zero is a
    # cost regression that no error path will ever surface.
    print(f"cache_hit_rate={hit_rate:.1%} read={u.cache_read_input_tokens} "
          f"write={u.cache_creation_input_tokens} uncached={u.input_tokens}")
    return response
 
# Warm the cache once, then measure the steady state.
cached_call("What is the escalation policy for a Tier 2 outage?")
cached_call("What is the refund window for annual plans?")

Run that twice in quick succession. The first call should show a non-zero write and a zero read; the second should show a large read and a near-zero write. If the second call still shows a write, something in your prefix is changing between requests — start by diffing the serialised request bodies byte for byte, because that is what the cache is comparing.

For per-prompt-version tracking over time rather than a one-off check, push these fields into your tracing layer. The AI Workshack guides to LangFuse and to LangSmith cover wiring cache and token fields into a trace so you can watch hit rate by prompt version, and the LiteLLM guide covers doing cost attribution at the gateway when you are caching across more than one provider.

What to do first

  1. Measure before you build. Log input tokens per request for a day and find the largest repeated prefix in your traffic. If no prefix repeats, caching has nothing to work with and you should be reading the article on controlling output tokens instead.
  2. Check that prefix against the minimum for every model you send it to — including any cheaper model you fall back to.
  3. Reorder the prompt: tools first and frozen, then the stable system content, then anything that varies. Move personalisation, timestamps and retrieved context after the breakpoint.
  4. Turn on the simplest form the provider offers — top-level `cache_control` on Anthropic, nothing at all on OpenAI, implicit on Google — and measure the hit rate for a week before reaching for explicit breakpoints or named cache objects.
  5. Only then consider the longer TTLs or Google's explicit caches, and only where you can show that reads will clear the higher write or storage cost.
Re-verify before you cite. Cache minimums, TTL options, multipliers and storage rates on this page were checked against Anthropic, OpenAI and Google documentation in August 2026 and all three providers change them without much ceremony. Treat the ratios as durable and the specific numbers as perishable.