In most RAG systems, retrieved context is over 80% of your input tokens and none of it caches. Here is how to retrieve less, pay less, and often answer better.

Prices in this guide are Claude API rates verified on 7 October 2026 and are used to make the arithmetic concrete. Model pricing changes — verify current rates before you budget. Third-party reranker and embedding prices are described directionally rather than quoted, because they vary by vendor and tier.

Why RAG quietly becomes your biggest token line item

Most teams tune a RAG pipeline for answer quality and never look at what it does to the bill. The reason it creeps up is structural: retrieval does not add a fixed cost to your application, it adds a multiplier to every single query.

A plain chat request sends a system prompt and a question. A RAG request sends a system prompt, a question, and every retrieved chunk — on every call. If you retrieve ten chunks of 800 tokens each, you have added 8,000 input tokens to each query, forever. Nothing is amortised. There is no cache that saves you, because the retrieved chunks change with each question.

That is why RAG cost scales with traffic in a way most people do not anticipate: the retrieved context, not the user's question and not the model's answer, usually dominates the input side of the invoice.

The cost formula

Every RAG query costs roughly:

input_tokens  = system_prompt + query + (top_k x avg_chunk_tokens) + conversation_history
output_tokens = answer length
 
cost = (input_tokens  x input_rate  / 1_000_000)
     + (output_tokens x output_rate / 1_000_000)

The term that matters is `top_k x avg_chunk_tokens`. It is the only one that multiplies, and it is the only one you fully control.

Here is a realistic baseline — a support-answering pipeline on Claude Sonnet 5.5 at $2 per million input tokens and $10 per million output tokens:

Component Tokens Share of input
System prompt 500 6%
User query 50 0.6%
Retrieved chunks (10 x 800) 8,000 93%
Total input 8,550 100%
Answer (output) 400 —

That is $0.0171 of input and $0.0040 of output, so about $0.0211 per query. At 100,000 queries a month, roughly $2,110 — of which about $1,975 is retrieved text. You are paying, overwhelmingly, to re-send your own knowledge base.

Before optimising anything, compute the retrieved-context share of your input tokens. If it is above about 80% — and in most RAG systems it is — retrieval tuning will beat prompt trimming by a wide margin, and it is the only place worth your attention first.

Lever 1: top-k is the biggest single win

`top_k` is the number of chunks you send to the model. It is usually set once during prototyping, often to a comfortable-feeling 8, 10 or 20, and then never revisited. Yet it scales your input cost linearly: halving top_k halves the retrieved-token cost.

The instinct against lowering it is reasonable — fewer chunks means a higher chance the answer is not in the context. But that instinct is usually calibrated against a weak retriever. If your retrieval ranks well, the answer is in the top few chunks and the rest are padding you pay for on every request.

The honest way to set it is empirically: measure how often the answer-bearing chunk appears within the top 1, 3, 5 and 10 results on a labelled question set. If recall at 3 is close to recall at 10, you are paying for seven chunks of insurance you do not need.

top_k (800-token chunks) Retrieved tokens Input cost / query Monthly at 100k queries
10 8,000 $0.0171 ~$2,110
5 4,000 $0.0091 ~$1,310
3 2,400 $0.0059 ~$990
Reranked to 3 from 30 candidates 2,400 $0.0059 + rerank ~$990 + rerank

Lever 2: reranking — retrieve wide, send narrow

This is the most useful structural idea in RAG cost control, and it rests on a price asymmetry: vector search and reranking are cheap, and LLM input tokens are not.

Instead of asking your vector store for the 10 chunks you will send, ask it for 30 or 50, run them through a reranker, and send the best 3. You improve precision and cut input tokens at the same time — the rare case where cost and quality move together.

The economics work because a reranker scores candidate passages at a tiny fraction of what a frontier model charges to read them. Retrieving 50 candidates costs essentially nothing in your vector store; reranking them costs a per-search fee measured in fractions of a cent on typical hosted rerankers; sending all 50 to the model would cost roughly 40,000 input tokens. Check your reranker's current rates, but the ordering almost always holds: search < rerank << LLM input.

# Retrieve wide, rerank, send narrow.
candidates = store.search(query, top_k=50)          # cheap
ranked     = reranker.rank(query, candidates)       # cheap, per-search fee
context    = ranked[:3]                             # only these become input tokens
 
prompt_context = "\n\n".join(c.text for c in context)
 
resp = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=600,
    system=[{
        "type": "text",
        "text": SYSTEM_PROMPT,
        "cache_control": {"type": "ephemeral"},   # stable prefix caches
    }],
    messages=[{"role": "user",
               "content": f"{prompt_context}\n\nQuestion: {query}"}],
)
print(resp.usage.input_tokens, resp.usage.output_tokens)
Reranking also lets you shrink chunks without losing recall. Small chunks retrieve precisely but individually carry less context; a reranker sorting 50 small chunks down to 4 often beats 10 large chunks on both answer quality and token count.

Lever 3: chunk size, and the waste it hides

Chunk size sets how much irrelevant text rides along with each relevant hit. A 1,500-token chunk retrieved because of one useful sentence means you paid for roughly 1,480 tokens of padding.

The tradeoff runs both ways, which is why there is no universal answer:

  • Smaller chunks (200-400 tokens): precise retrieval, little waste per hit, but each chunk carries less surrounding context, so you often need more of them — and you risk splitting an answer across a boundary.
  • Larger chunks (1,000-1,500 tokens): more self-contained and robust to boundary splits, but every hit drags in padding you pay for on every query.
  • Overlap multiplies storage and retrieval duplication; heavy overlap means the same sentences can appear in two retrieved chunks, so you pay twice for the same text in one request.

A practical default for prose documentation is moderate chunks (around 400-800 tokens) with modest overlap, paired with reranking to recover the precision that larger chunks lose. For reference material with hard structural boundaries — API endpoints, FAQ entries, product records — chunk on the structure instead of a token count, because the natural unit is already the right unit.

Lever 4: metadata filtering before you ever embed a query

The cheapest chunk is one your retriever never considers. If a query is scoped to one tenant, one product version, or one document category, filter on metadata before the similarity search rather than retrieving broadly and hoping the ranking sorts it out.

This cuts cost indirectly but reliably: a smaller, more relevant candidate pool means a higher proportion of your top_k is actually useful, which is what lets you lower top_k without losing recall. It also removes a whole class of wrong-tenant and wrong-version answers.

What prompt caching can and cannot do for RAG

This is where a lot of RAG cost advice goes wrong, so it is worth being precise.

Prompt caching works on a stable prefix. Your system prompt, your tool definitions and any fixed instructions are stable, so they cache well and a cache read costs a fraction of standard input — currently 0.1x base input on most models, and lower still on the newest ones (0.05x on Claude Opus 5.5 and Sonnet 5.5). Caching those is close to free money.

Your retrieved chunks are not stable. They change with every question, they sit after the cached prefix, and they will not be served from cache. So caching does not reduce the cost of retrieved context — the dominant term — at all.

If you have read that "prompt caching makes RAG cheap", treat it sceptically. Caching makes the fixed part of your prompt cheap. The retrieved part, which is usually over 80% of your input tokens, is different on every request and is billed at full rate every time. The exception is a genuinely static corpus small enough to sit in the prompt — but at that point you are not doing retrieval, you are doing long context, which is a different tradeoff.

Measuring it, so you do not cut blind

Every lever here trades tokens against the risk of missing the answer. Cutting top_k without measuring recall is how a cost optimisation quietly becomes a quality regression that shows up weeks later as complaints rather than as a metric.

Track two things per query and watch them together:

  • Cost: input tokens, output tokens, and the retrieved-context share of input. Read these from the response `usage` object rather than estimating.
  • Quality: whether the answer-bearing chunk was retrieved at your chosen top_k (recall@k), and whether the final answer was grounded in it. A labelled question set of even 50-100 examples is enough to detect a regression.

Run the before-and-after on the same question set. A change that cuts tokens 45% and recall by one percentage point is a clear win; one that cuts tokens 45% and recall by fifteen points is a quality cut dressed up as a saving.

# Per-query accounting you can log and aggregate.
ctx_tokens = client.messages.count_tokens(
    model="claude-sonnet-5-5",
    messages=[{"role": "user", "content": prompt_context}],
).input_tokens
 
log.info({
    "input_tokens":   resp.usage.input_tokens,
    "output_tokens":  resp.usage.output_tokens,
    "cache_read":     resp.usage.cache_read_input_tokens,
    "retrieved_tokens": ctx_tokens,
    "retrieved_share": ctx_tokens / resp.usage.input_tokens,
    "top_k": len(context),
})

For the evaluation side of this — building the question set, scoring groundedness, and wiring it into CI so a retrieval change cannot regress quality unnoticed — see our RAG evaluation guides, which cover Ragas and the wider eval tooling in detail.

What this costs you

None of these levers is free, and it is worth being plain about the price of each:

Lever Typical saving What it costs you
Lower top_k Linear in the reduction Recall risk — the answer may not be in context. Must be measured.
Reranking 50-70% of retrieved tokens A reranker fee per search, extra latency in the retrieval hop, and another dependency to run.
Smaller chunks Less padding per hit More boundary splits; usually needs reranking to hold quality.
Metadata filtering Indirect — enables lower top_k Metadata discipline at ingest; stale or missing metadata silently hides documents.
Caching the system prompt 0.1x-0.05x on the fixed prefix only Nothing meaningful, but it does not touch retrieved context.

A tuning sequence that works

  1. Instrument first. Log input tokens, output tokens, and the retrieved-context share per query. You cannot prioritise without knowing whether retrieval is 40% or 93% of your input.
  2. Build a small labelled question set — 50-100 questions with known answer locations. This is your regression net for everything that follows.
  3. Measure recall at k = 1, 3, 5, 10 on that set. This tells you how much insurance your current top_k is buying.
  4. Add a reranker, retrieve wide (30-50 candidates), and send the top 3-5. Re-measure recall; it usually rises while tokens fall.
  5. Cache the system prompt and any fixed instructions. Confirm with `cache_read_input_tokens` that hits are actually happening.
  6. Add metadata filtering where queries have a natural scope, then re-test whether top_k can go lower still.
  7. Re-run the question set and compare cost and recall side by side before shipping.

Quick reference

  • Retrieved context is usually over 80% of RAG input tokens — tune retrieval before you trim prompts.
  • top_k multiplies cost linearly and is the fastest lever available.
  • Retrieve wide, rerank, send narrow: search and reranking are far cheaper than frontier-model input tokens.
  • Prompt caching helps the fixed prefix, not the retrieved chunks — they differ every request.
  • Heavy chunk overlap can make you pay twice for the same sentences in one request.
  • Never cut top_k without a recall measurement on a labelled set.
  • Read `usage` from the response rather than estimating token counts.