A long conversation is not expensive because it is long. It is expensive because you re-send all of it, every single turn.
There is a specific moment in every chat product's life where someone opens the billing dashboard and finds that a small number of users are responsible for a large share of the spend. They are not sending more messages than everyone else. They are sending them in one long thread.
This article is about why that happens, and about the four mechanisms — truncation, rolling summarisation, server-side clearing, and server-side compaction — that fix it. Each one buys you tokens and costs you something else. The goal is not to pick a favourite; it is to know which cost you are choosing to pay.
All prices, multipliers, minimums and API parameters in this article were verified against provider documentation as of August 2026. Model pricing and cache mechanics change frequently — verify current pricing and limits against the provider's own docs before you build a business case on them.The mechanism: chat APIs are stateless, so history is an input
The Messages API and its equivalents at other providers hold no server-side memory of your conversation. Every request is self-contained. When you send turn ten, you are sending turns one through nine as well, because that is the only way the model can see them. The provider bills those re-sent tokens as input tokens on every request — not once, but every time.
That single fact turns a linear-looking product into a quadratic-looking bill. If each exchange adds roughly the same number of tokens, then turn N costs about N units of input, and the cumulative cost across N turns is proportional to N squared over two. A conversation that is twice as long does not cost twice as much. It costs roughly four times as much.
Here is what that looks like with a 2,000-token system prompt and roughly 600 tokens added per exchange. The numbers are illustrative, but the shape is not — it is arithmetic, not an estimate.
| Turn | History sent this turn | Cumulative input tokens billed |
|---|---|---|
| 1 | 2,600 | 2,600 |
| 5 | 5,000 | 19,000 |
| 10 | 8,000 | 52,000 |
| 25 | 17,000 | 247,000 |
| 50 | 32,000 | 857,000 |
| 100 | 62,000 | 3,232,000 |
Turn 100 is only about 24 times bigger than turn 1, but the running total is more than 1,200 times bigger. Most teams instrument the per-request number and never plot the cumulative one, which is why the discovery usually arrives via the invoice.
Measure before you cut
Do not start by guessing a window size. Start by finding out what your actual conversations look like. Two measurements are enough.
The first is a pre-flight count: how many tokens is this request, before you send it? Anthropic exposes a dedicated endpoint for this, and it counts the whole rendered request — system prompt, tools and messages — not just the text you can see.
import anthropic
client = anthropic.Anthropic()
def conversation_size(system: str, messages: list) -> int:
"""Exact token count for a request, before you pay for it."""
return client.messages.count_tokens(
model="claude-sonnet-5",
system=system,
messages=messages,
).input_tokens
# Plot this across the life of a real conversation, not a synthetic one.
for turn in range(1, len(messages), 2):
print(turn, conversation_size(SYSTEM_PROMPT, messages[:turn]))The second measurement is what you were actually billed, which lives in the usage object on every response. On Anthropic this is three fields, and the total prompt size is their sum — a detail that trips people up constantly. If your agent ran for an hour and input_tokens reads 4,000, the rest was served from cache; check the sum, not the single field.
u = response.usage
total_prompt = (
u.input_tokens # uncached, billed at full rate
+ u.cache_creation_input_tokens # written to cache this request
+ u.cache_read_input_tokens # served from cache
)Do not estimate Claude token counts with tiktoken. It is OpenAI's tokenizer and it undercounts Claude tokens materially on ordinary prose, and by much more on code. Separately: Claude models from 4.7 onward use a newer tokenizer that produces roughly 30% more tokens for the same text than earlier models, so a budget calibrated on an older model will be wrong on a newer one. Re-baseline when you change models.Four ways to stop paying for the whole history
Every technique in this space is a variation on one of four moves: drop old turns, replace old turns with a summary, have the provider drop them for you, or have the provider summarise them for you. They differ in who does the work, what gets lost, and what happens to your prompt cache.
| Technique | What it does | Token effect | What it costs you |
|---|---|---|---|
| Sliding window (truncate) | Keep the last N turns or N tokens | Caps per-request input at a constant | Hard forgetting — dropped facts are gone with no trace |
| Rolling summarisation | Replace old turns with a generated summary | Compresses history to a small constant | An extra model call, latency spikes, and lossy/drifting summaries |
| Server-side context editing | Provider clears old tool results in place | Removes the bulkiest agent content | Cache invalidation on every clear; results are gone, not summarised |
| Server-side compaction | Provider summarises history near a threshold | Keeps you inside the window automatically | Least control over what survives; beta surface |
Sliding windows: the cheapest thing that works
A sliding window keeps the system prompt and the most recent slice of the conversation, and discards the middle. It is trivially cheap to implement, adds zero latency, and caps per-request cost at a constant instead of letting it grow. For a large share of assistants — anything where turn 30 rarely depends on turn 3 — it is the correct answer and nothing more sophisticated is warranted.
LangChain ships a utility for this that handles the awkward parts: keeping the system message, and making sure the trimmed list still starts on a human turn so the API does not reject it.
from langchain_core.messages import trim_messages
from langchain.chat_models import init_chat_model
model = init_chat_model("anthropic:claude-sonnet-5")
trimmed = trim_messages(
messages,
max_tokens=4000,
strategy="last", # keep the most recent messages
token_counter=model, # count with the model that will be billed
include_system=True, # never drop the system prompt
start_on="human", # keep the sequence API-valid
allow_partial=False, # do not cut a message in half
)The tradeoff is blunt and worth stating plainly: truncation causes hard forgetting. The user's name, the constraint they gave in turn two, the file they said not to touch — all gone, with no signal to the model that anything is missing. The failure mode is not an error; it is an assistant that confidently contradicts something the user told it twenty minutes ago.
Rolling summarisation: pay a little to forget less
Summarisation replaces the dropped turns with a compressed account of them. When the history crosses a threshold, you make an extra model call that produces a summary, then continue with the summary plus a short tail of recent turns.
LangChain's agent middleware implements this pattern directly, with the trigger and retention policy as configuration rather than code you own.
from langchain.agents import create_agent
from langchain.agents.middleware import SummarizationMiddleware
agent = create_agent(
model="anthropic:claude-sonnet-5",
tools=[...],
middleware=[
SummarizationMiddleware(
# Summarise with a cheaper model than the one answering.
model="anthropic:claude-haiku-4-5",
trigger=("tokens", 8000), # compress once history crosses this
keep=("messages", 20), # verbatim tail kept after the summary
)
],
checkpointer=checkpointer,
)Summarisation is not free, and the ways it costs you are easy to under-model. Each compaction is a full extra request whose input is the history you are about to throw away — so the compression itself is billed at roughly the size of the thing being compressed. It lands as a latency spike on whichever unlucky user's turn triggers it. And it is lossy in a compounding way: summarising a summary loses detail on every pass, so a very long conversation ends up with an increasingly vague account of its own early history.
Summarise with a cheaper model than the one answering. Compression is a well-specified, low-judgement task — exactly the shape that a small fast model handles well. Running summarisation on your flagship model is one of the most common avoidable line items in this whole area.Where the summary goes — and why it is a cost decision, not a layout one
The obvious place to put a rolling summary is inside the system prompt. It reads as context about the conversation, so it feels like system-prompt material. This is the single most expensive mistake in this article.
Prompt caching is a prefix match. The provider hashes your request from the beginning, and a cache entry only hits if every byte up to the breakpoint is identical. The render order on Anthropic is tools, then system, then messages. Putting a summary that changes every few turns inside the system prompt means the prefix changes every few turns, which means the cached tools and cached system prompt behind it are invalidated too. You save tokens on the history and lose the 0.1x multiplier on everything ahead of it.
Put the summary after the stable prefix instead. Three placements work, in descending order of preference.
- As the first message in the messages array — a user or assistant turn containing the summary, followed by the retained recent turns. Portable across every provider, and it leaves the tools and system prompt cacheable.
- As a mid-conversation system message — appending {"role": "system", "content": ...} to messages[] rather than editing top-level system. This carries operator authority without touching the cached prefix. Available on Claude Opus 5, Claude Opus 4.8, Claude Fable 5 and Claude Mythos 5; it returns a 400 on models that do not support it, so catch that and fall back.
- Inside the top-level system prompt — only if you are not caching at all, or if the summary genuinely updates less often than your cache TTL.
The same logic explains a less obvious rule: trim rarely, and in large chunks. Every trim changes the front of the message history, which invalidates the cached message prefix. Trimming one message per turn means you invalidate the cache every turn and never get a read. Trimming forty messages once every forty turns costs you one invalidation and leaves the cache productive in between. Chunky, infrequent trimming beats smooth, continuous trimming on cost even though it looks worse on a graph of context size.
Letting the provider do it
Two server-side mechanisms handle this without you maintaining a message list at all. They are worth knowing because agentic workloads — where the bulk of the history is tool results, not conversation — respond much better to them than to generic truncation.
Context editing: clear old tool results
In an agent loop, the largest thing in the context is usually not what anyone said. It is the output of a file read, a search, or an API call from fifteen steps ago that nobody will ever need again. Context editing clears those in place, on the server, based on a threshold you set.
response = client.beta.messages.create(
model="claude-opus-5",
max_tokens=8192,
betas=["context-management-2025-06-27"],
tools=[...],
messages=messages,
context_management={
"edits": [
{
"type": "clear_tool_uses_20250919",
"trigger": {"type": "input_tokens", "value": 30000},
"keep": {"type": "tool_uses", "value": 3},
# Clear a meaningful amount when you do clear, so the
# cache invalidation is amortised over real savings.
"clear_at_least": {"type": "input_tokens", "value": 5000},
"exclude_tools": ["web_search"],
}
]
},
)
for edit in response.context_management.applied_edits:
print(edit.type, edit.cleared_input_tokens)Note clear_at_least. Clearing invalidates the cached prefix and triggers a cache write on the next request, so a clear that removes 400 tokens is a net loss. That parameter is the guard that stops the feature from firing on trivially small wins. The defaults, if you set nothing, are a 100,000-token trigger keeping the last three tool uses.
Compaction: server-side rolling summarisation
Compaction is the managed version of the summarisation pattern above. When the conversation approaches a configurable threshold — 150,000 input tokens by default, with a floor of 50,000 — the API summarises earlier context itself and returns a compaction block.
response = client.beta.messages.create(
betas=["compact-2026-01-12"],
model="claude-opus-5",
max_tokens=8192,
messages=messages,
context_management={
"edits": [
{"type": "compact_20260112",
"trigger": {"type": "input_tokens", "value": 150000}}
]
},
)
# CRITICAL: append the whole content list, not just the text.
# The compaction block is how the API knows what it already summarised.
messages.append({"role": "assistant", "content": response.content})The most common way to break compaction is to extract the text from the response and append only that. The compaction block must be passed back verbatim; without it the API cannot replace the compacted history and you silently lose the state. Append response.content, never a plain string.How to actually choose a window size
Window size is usually picked by feel, then never revisited. It is a budget decision and it can be derived. Work backwards from what you are willing to spend on a conversation.
- Decide your per-conversation input budget — a hard number, from your unit economics, not a vibe.
- Measure your real distribution: median and 90th-percentile turns per conversation, and median tokens added per exchange. Take these from production traffic, not from your own testing.
- Compute the fixed overhead you pay every turn regardless: system prompt plus tool schemas. On a long conversation this is often larger than the history you are worrying about — which is why the companion article on system-prompt and tool-schema bloat is frequently the higher-leverage fix.
- Set the window so that (fixed overhead + window) multiplied by your 90th-percentile turn count lands inside the budget.
- Then check the window against your prompt-cache minimum. Trimming a prompt below the model's minimum cacheable prefix silently disables caching — no error, just a permanent 10x jump on the portion that used to be a cache read.
That last point deserves emphasis because it is genuinely counterintuitive. Making a prompt shorter can make it more expensive. The minimum cacheable prefix is model-specific and not monotonic across generations — it is 1,024 tokens on Claude Opus 4.8 and Claude Sonnet 5, but 4,096 tokens on Claude Haiku 4.5. A 3,000-token prompt caches on Sonnet 5 and silently does not cache on Haiku 4.5. Check the number for every model you route to, including cheap fallbacks.
Measuring the quality cost of forgetting
Every technique here degrades memory. The question is not whether, but how much and where — and you cannot answer that by reading transcripts, because the failures are rare, specific, and invisible in aggregate metrics. You need a probe set.
The construction is simple and takes an afternoon. Take twenty to fifty real long conversations. For each, write a question whose answer depends on something stated early and never repeated: a constraint, a name, a preference, a decision. Then run each conversation through each candidate policy and score whether the assistant still gets it right.
| Policy | Avg input tokens/turn | Relative cost | Early-fact recall |
|---|---|---|---|
| No management (full history) | Grows without bound | Baseline (highest) | Ceiling — the reference score |
| Sliding window, small | Constant, low | Lowest | Drops sharply on long threads |
| Sliding window, large | Constant, moderate | Low | Holds until the window is exceeded |
| Rolling summarisation | Constant, low | Low + per-compaction call | Degrades gradually rather than falling off a cliff |
| Context editing (agents) | Constant, moderate | Low + cache re-writes | Unaffected for conversation; tool history is gone |
Fill that table with your own numbers and the decision makes itself. The pattern most teams find is that recall is nearly flat as the window shrinks, right up until it collapses — and the useful window sits just above that cliff. You cannot find the cliff without measuring for it. If you already run an evaluation harness for retrieval, this is the same discipline applied to memory; the AI Workshack RAG Evaluation track covers building and scoring that harness in detail.
The honest tradeoff summary
- Truncation costs you quality, silently. It is free in latency and complexity, and it will eventually make your assistant contradict the user. Mitigate with a large window and a probe set, not with hope.
- Summarisation costs you latency and money at the moment it fires. It converts a quality cliff into a quality slope, which is usually the better failure mode, but it adds a second model call and a second thing to debug.
- Both of them cost you cache hits if you do them badly. Trim in big infrequent chunks, keep the summary out of the system prompt, and treat the cached prefix as the thing you are protecting.
- Server-side options cost you control. Context editing and compaction are less code and less thinking, but you choose less about what survives — and both are beta surfaces whose parameters can move.
What to do first
- Plot cumulative input tokens per conversation for a week of real traffic. If the curve is not visibly quadratic, this article is not your problem and you should look at output tokens or system-prompt overhead instead.
- Measure your fixed per-call overhead — system prompt plus tool schemas — before touching history. If it dominates, fix that first; it is a bigger and easier win.
- Turn on prompt caching and confirm you are getting reads, before you optimise history. Caching and trimming interact, and tuning one blind to the other wastes effort.
- Implement a sliding window with a deliberately generous size. Ship it. It is the cheapest thing that works and it caps your worst case immediately.
- Build the twenty-question probe set and find your recall cliff. Only then decide whether the extra call for summarisation is buying you anything you can measure.
Related reading on AI Workshack
- System-Prompt and Tool-Schema Bloat: The Tax You Pay on Every Call — the fixed overhead that sits underneath every number in this article.
- Prompt Caching Across Anthropic, OpenAI and Google: What Each One Actually Does — the multipliers, minimums and silent invalidators referenced throughout.
- Why AI Agents Burn Tokens: The Agent Loop, Re-Sent Context, and Where the Spend Actually Goes — the agentic version of the same quadratic problem.
- LangGraph Memory Explained: Checkpoints, Threads and Stores — and LangGraph Long-Term Memory Using Store — for persisting the facts you deliberately trim out of the window.
- The RAG Evaluation track — for building the scored harness that turns "it feels fine" into a number.