Cascade patterns, routing heuristics that do not need an LLM, the break-even maths on escalation, and how to prove you did not quietly degrade quality.
The advice is always the same: use a cheaper model for simple tasks. It is correct and almost useless, because it omits every hard part. Which tasks are simple? How do you find out without shipping a regression? What does a cascade actually cost when it escalates? And why did the saving come in smaller than the price sheet predicted?
This article works through model tiering as an engineering problem: the ratios that matter, three routing patterns with their real cost structures, the confidence signals that are worth acting on, and the measurement discipline that separates a genuine saving from an undetected quality regression.
Prices and model capabilities cited here were verified against Anthropic's documentation as of August 2026. Provider pricing changes frequently — verify current pricing before you build a business case. Where possible this article uses ratios, which decay much more slowly than absolute figures.Work in ratios, not dollars
Absolute prices are the most perishable thing in this field. The ratio between tiers is far more stable, and it is the only number a routing decision actually needs.
| Model | Input / MTok | Output / MTok | Relative to the cheapest tier |
|---|---|---|---|
| Claude Opus 4.8 (`claude-opus-5`) | $5 | $25 | 5x |
| Claude Sonnet 5 (`claude-sonnet-5`) | $2 | $10 | 2x |
| Claude Haiku 4.5 (`claude-haiku-4-5`) | $1 | $5 | 1x |
Two things fall out of that table. The output-to-input ratio is a consistent 5:1 across all three tiers, so moving down a tier does not change the shape of your bill — it scales it. And the full spread from top tier to bottom is 5x, not 50x. That bounds the prize: even a perfect routing system that pushed every request to the cheapest model would cut inference cost by 80%, and a realistic one that moves half the traffic saves far less.
Before building any routing layer, compute the ceiling. Take your current spend, assume every request that could plausibly move down a tier does, and calculate the saving. If that number is smaller than a week of engineering time, stop — go and read the article on controlling output tokens instead, which usually has a better effort-to-saving ratio.The cheapest tier change is not a model swap
Before splitting traffic across models, use the cost dial that lives inside a single model. On Claude Opus 4.8 and Claude Sonnet 5, `output_config.effort` controls how much the model thinks and how many tokens it spends overall. The default is `high`. Dropping it costs you one parameter and no routing infrastructure.
import anthropic
client = anthropic.Anthropic()
# Same model, same prompt, materially less token spend.
response = client.messages.create(
model="claude-sonnet-5",
max_tokens=2048,
thinking={"type": "adaptive"},
output_config={"effort": "low"}, # low | medium | high | xhigh | max; default is high
messages=[{"role": "user", "content": task}],
)
print(response.usage.input_tokens, response.usage.output_tokens)Sweep effort across your eval set before you sweep models. It is one line of code, it keeps a single prompt and a single eval suite, and it leaves your prompt cache intact. Note that Claude Haiku 4.5 does not support the effort parameter — the dial exists on the higher tiers, which is convenient, because those are the tiers where it saves the most.
Pattern 1: static routing on deterministic signals
The cheapest router is one that never calls a model. Route on properties of the request you already know, before any inference happens.
- **Task type.** You usually know this from the endpoint. Classification, extraction, tagging, routing and short rewrites go to the cheap tier. Multi-step reasoning, code generation and long-form synthesis go up.
- **Input size.** Use the free token-counting endpoint to measure before you route. Note that Claude Haiku 4.5 has a 200K context window against 1M on Opus 4.8 and Sonnet 5, so this is a hard constraint and not just a heuristic.
- **Presence of tools or images.** Tool-heavy and multimodal work is where cheap models degrade first. This is a strong and cheap signal.
- **Expected output length.** Haiku 4.5 caps at 64K output tokens against 128K on the higher tiers.
- **Customer or plan tier.** A defensible business rule rather than a technical one, and often the highest-leverage router in a SaaS product.
import anthropic
client = anthropic.Anthropic()
CHEAP = "claude-haiku-4-5"
MID = "claude-sonnet-5"
TOP = "claude-opus-5"
def choose_model(task_type: str, messages, system: str, tools=None) -> str:
if task_type in {"classify", "extract", "tag", "route"} and not tools:
# count_tokens is free and rate-limited separately from message creation.
# Recount per target model: token counts are NOT comparable across
# Claude tokenizer generations.
count = client.messages.count_tokens(
model=CHEAP, system=system, messages=messages
).input_tokens
if count < 150_000: # leave headroom inside Haiku's 200K window
return CHEAP
return MID
if task_type in {"reason", "plan", "generate_code"}:
return TOP
return MIDDo not reuse a token count measured on one model to make a decision about another. Anthropic's documentation notes that Claude 4.7 and later models use a newer tokenizer that produces roughly 30% more tokens for the same text, and advises recounting against the model you plan to use. A routing rule calibrated on one tokenizer generation will misjudge context fit on another.Pattern 2: cascade — cheap first, escalate on failure
A cascade sends everything to the cheap model, checks the result, and reruns on an expensive model when the check fails. It is more accurate than static routing because the decision is made after seeing an attempt rather than before. It is also the pattern people most often get wrong economically, because escalated requests pay twice.
The break-even maths
Let C_cheap be the cost of the cheap attempt, C_exp the cost of the expensive one, and r the fraction of requests that escalate. Always-expensive costs C_exp per request. The cascade costs C_cheap + r x C_exp. The cascade wins when:
C_cheap + (r x C_exp) < C_exp
=> r < 1 - (C_cheap / C_exp)
Haiku 4.5 -> Opus 4.8, same token counts, cost ratio 1:5:
r < 1 - 0.2 = 0.80
Sonnet 5 -> Opus 4.8, cost ratio 2:5:
r < 1 - 0.4 = 0.60An 80% escalation ceiling sounds like enormous headroom, and on tokens alone it is. But the money break-even is not the decision break-even, because escalation costs more than money:
- **Latency is additive on escalation.** An escalated request waits for the cheap model to fail before the expensive one starts. At a 30% escalation rate, roughly a third of your users see something close to double the latency. That is a p90 regression that will not show up in your average.
- **Your prompt cache fragments.** Each model keeps its own cache. Splitting traffic 70/30 means each model's prefix is warmed less often, and with a five-minute TTL a low-traffic tier can drop below its hit threshold entirely. In a lopsided split, the minority model may write its cache on nearly every request at a 1.25x multiplier and read it almost never.
- **The minimum cacheable prefix differs by tier.** Claude Haiku 4.5 needs 4,096 tokens to cache; Opus 4.8 and Sonnet 5 need 1,024. A shared 2,000-token system prompt caches on the higher tiers and silently does not cache on Haiku.
- **Tool overhead is not uniform.** The tool-use system prompt that the API injects when tools are present costs a different number of tokens on each model. On the figures published in August 2026, it is smaller on Opus 4.8 than on Haiku 4.5 — so a tool-heavy request narrows the tier gap slightly rather than widening it.
A practical planning rule: treat escalation rates above about 30% as a signal that the cheap tier is the wrong model for this task, regardless of what the token arithmetic says. At that point you are paying two latencies and running two eval suites to save a fraction you could get more cleanly by lowering effort on a single model.Confidence signals that actually work
The cascade lives or dies on the escalation check. The single worst check is asking the model how confident it is in prose — self-reported confidence from a small model is poorly calibrated and tends to be uniformly high. Prefer signals that are mechanical and verifiable outside the model.
| Signal | How you get it | Quality |
|---|---|---|
| Downstream validation failure | Does the SQL parse? Does the code compile? Does the ID exist? | Best — objective and free |
| Schema violation | Structured output rejected, or a required field missing | Excellent |
| Truncation | `stop_reason == "max_tokens"` | Excellent — a clear signal the task overflowed |
| Explicit abstention | A schema field the model can set to decline, e.g. `insufficient_information` | Good — you are asking for a decision, not a feeling |
| Enum-constrained confidence | A `confidence` field restricted to `high`/`medium`/`low` in a strict schema | Fair — calibrate it against ground truth before trusting it |
| Retrieval score | Similarity of the supporting context you passed in | Good, and available before you call the model at all |
| "Are you sure?" in prose | Free-text self-assessment | Poor — do not build on it |
import anthropic
from pydantic import BaseModel
from typing import Literal
client = anthropic.Anthropic()
class Triage(BaseModel):
category: Literal["billing", "technical", "account", "other"]
# An explicit abstention channel. The model is deciding, not introspecting.
sufficient_information: bool
def cascaded_triage(ticket: str) -> tuple[Triage, str]:
cheap = client.messages.parse(
model="claude-haiku-4-5",
max_tokens=256,
messages=[{"role": "user", "content": ticket}],
output_format=Triage,
)
result = cheap.parsed_output
# Escalate on abstention or on truncation — both are mechanical checks.
if result.sufficient_information and cheap.stop_reason != "max_tokens":
return result, "claude-haiku-4-5"
strong = client.messages.parse(
model="claude-opus-5",
max_tokens=1024,
messages=[{"role": "user", "content": ticket}],
output_format=Triage,
)
return strong.parsed_output, "claude-opus-5"Return the model that served the request, and log it. Escalation rate is the single most important operational metric a cascade has — it is both your cost driver and your earliest warning that input distribution has shifted.
Pattern 3: an LLM router
You can also ask a cheap model to classify the request and pick a tier. This is the most flexible pattern and the one that most often fails to pay for itself, because the router call sits on the critical path of every single request — including the ones that were always going to the cheap tier anyway.
It is worth it only when three things hold at once: the router is dramatically cheaper than the average request it routes (a 200-token classification against a 20,000-token task, not against a 500-token one), the routing decision genuinely cannot be made from static signals, and the added latency is acceptable on every request rather than just escalated ones. If you are considering an LLM router, first check whether a static rule on task type and input size gets you most of the way there. It usually does.
Implementing this without writing a router
Everything above is routing policy. The plumbing — a unified interface across providers, fallback chains, retries, per-model rate limits, per-key budgets and cost accounting — is a solved problem, and writing it yourself is rarely a good use of time. The AI Workshack guide to LiteLLM covers standing up the proxy with fallbacks and cost tracking, and the follow-up on running the LiteLLM proxy in production covers load balancing, caching and team access. Put your policy in that layer rather than scattering model names through application code.
The practical benefit is that model choice becomes configuration. When a new tier ships, you change a routing rule and re-run your evals rather than grepping for model IDs across a codebase.
Measuring the quality loss (this is the actual work)
Routing without measurement is not cost optimisation, it is a silent quality cut with a cost saving attached. The saving is visible on the invoice within days; the regression shows up in churn several months later. Build the measurement first.
- **Freeze an eval set from real traffic.** A few hundred real inputs, sampled to reflect your actual distribution — not a hand-picked set of easy cases. Include the awkward ones you would rather not look at.
- **Establish a reference.** Run the expensive model over the set and treat its outputs as the baseline. You are measuring the delta from what you ship today, not from an abstract notion of correctness.
- **Score agreement, not vibes.** For classification, exact match. For extraction, field-level accuracy. For open text, a judge model with a rubric, validated against a human-labelled subset before you trust it. Write down the metric before you run the experiment.
- **Break results down by segment.** Aggregate accuracy hides everything that matters. A cheap model that scores 94% overall may be at 99% on your biggest customer type and 61% on a smaller one. The average will not tell you, and the smaller segment is often the one that complains loudest.
- **Shadow-run before switching.** Send production traffic to both models, serve the expensive result, log both. You get real-distribution comparison data with zero user risk. It costs you double for the duration — budget for it as an experiment.
- **Set a quality floor and treat it as a release gate.** Decide in advance what agreement rate is acceptable per segment, and hold the routing change to it exactly as you would hold a code change to a test suite.
For the trace-and-score infrastructure — capturing model, prompt version, token counts and eval scores per request so you can slice them later — the AI Workshack guides to LangFuse and to the LangSmith comparison cover the setup. The key requirement is that model identity is a first-class dimension in your traces, so "which model served this" is a filter rather than an archaeology project.
Watch for these after you ship
- **Escalation rate drift.** A cascade that escalated 20% of requests in testing and 55% in production means your eval set did not match reality. It is also costing you more than the single-model baseline.
- **Cache hit rate falling on both tiers.** The classic routing side effect. Check it explicitly after every routing change.
- **Prompt divergence.** Two models, two prompts, and over a few months only one of them gets maintained. Version the prompts together, and re-run both eval suites on every prompt change.
- **Quality complaints that do not map to errors.** The cheap tier rarely fails loudly. It fails by being slightly worse — a little less specific, a little more generic. That shows up in user sentiment long before it shows up in a metric, so keep a qualitative review loop.
The honest summary
| Approach | Typical saving | What it costs you |
|---|---|---|
| Lower `effort` on one model | Meaningful, workload-dependent | Some reasoning depth. One parameter, one eval suite, cache intact. |
| Static routing on task type | Up to the tier ratio on routed traffic | Two prompts, two eval suites, split cache. |
| Cascade with escalation | Good below ~30% escalation | Latency on escalated requests, double cost on those requests, most complexity. |
| LLM router | Situational | A model call on every request's critical path. Usually not worth it. |
Start at the top of that table and work down. Effort tuning is nearly free and reversible. Static routing is a day of work. A cascade is a system with its own failure modes and its own operational metric. Each step down buys a bit more saving for a lot more complexity, and most teams should stop before the bottom row.
Model IDs, tier prices, context windows and parameter support in this article were verified against Anthropic documentation in August 2026. New tiers ship regularly and pricing moves — re-check the current lineup and re-run your evals before rolling a routing change forward.