Cascade patterns, routing heuristics that do not need an LLM, the break-even maths on escalation, and how to prove you did not quietly degrade quality.

The advice is always the same: use a cheaper model for simple tasks. It is correct and almost useless, because it omits every hard part. Which tasks are simple? How do you find out without shipping a regression? What does a cascade actually cost when it escalates? And why did the saving come in smaller than the price sheet predicted?

This article works through model tiering as an engineering problem: the ratios that matter, three routing patterns with their real cost structures, the confidence signals that are worth acting on, and the measurement discipline that separates a genuine saving from an undetected quality regression.

Prices and model capabilities cited here were verified against Anthropic's documentation as of August 2026. Provider pricing changes frequently — verify current pricing before you build a business case. Where possible this article uses ratios, which decay much more slowly than absolute figures.

Work in ratios, not dollars

Absolute prices are the most perishable thing in this field. The ratio between tiers is far more stable, and it is the only number a routing decision actually needs.

Model Input / MTok Output / MTok Relative to the cheapest tier
Claude Opus 4.8 (`claude-opus-5`) $5 $25 5x
Claude Sonnet 5 (`claude-sonnet-5`) $2 $10 2x
Claude Haiku 4.5 (`claude-haiku-4-5`) $1 $5 1x

Two things fall out of that table. The output-to-input ratio is a consistent 5:1 across all three tiers, so moving down a tier does not change the shape of your bill — it scales it. And the full spread from top tier to bottom is 5x, not 50x. That bounds the prize: even a perfect routing system that pushed every request to the cheapest model would cut inference cost by 80%, and a realistic one that moves half the traffic saves far less.

Before building any routing layer, compute the ceiling. Take your current spend, assume every request that could plausibly move down a tier does, and calculate the saving. If that number is smaller than a week of engineering time, stop — go and read the article on controlling output tokens instead, which usually has a better effort-to-saving ratio.

The cheapest tier change is not a model swap

Before splitting traffic across models, use the cost dial that lives inside a single model. On Claude Opus 4.8 and Claude Sonnet 5, `output_config.effort` controls how much the model thinks and how many tokens it spends overall. The default is `high`. Dropping it costs you one parameter and no routing infrastructure.

import anthropic
 
client = anthropic.Anthropic()
 
# Same model, same prompt, materially less token spend.
response = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=2048,
    thinking={"type": "adaptive"},
    output_config={"effort": "low"},  # low | medium | high | xhigh | max; default is high
    messages=[{"role": "user", "content": task}],
)
 
print(response.usage.input_tokens, response.usage.output_tokens)

Sweep effort across your eval set before you sweep models. It is one line of code, it keeps a single prompt and a single eval suite, and it leaves your prompt cache intact. Note that Claude Haiku 4.5 does not support the effort parameter — the dial exists on the higher tiers, which is convenient, because those are the tiers where it saves the most.

Pattern 1: static routing on deterministic signals

The cheapest router is one that never calls a model. Route on properties of the request you already know, before any inference happens.

  • **Task type.** You usually know this from the endpoint. Classification, extraction, tagging, routing and short rewrites go to the cheap tier. Multi-step reasoning, code generation and long-form synthesis go up.
  • **Input size.** Use the free token-counting endpoint to measure before you route. Note that Claude Haiku 4.5 has a 200K context window against 1M on Opus 4.8 and Sonnet 5, so this is a hard constraint and not just a heuristic.
  • **Presence of tools or images.** Tool-heavy and multimodal work is where cheap models degrade first. This is a strong and cheap signal.
  • **Expected output length.** Haiku 4.5 caps at 64K output tokens against 128K on the higher tiers.
  • **Customer or plan tier.** A defensible business rule rather than a technical one, and often the highest-leverage router in a SaaS product.
import anthropic
 
client = anthropic.Anthropic()
 
CHEAP = "claude-haiku-4-5"
MID = "claude-sonnet-5"
TOP = "claude-opus-5"
 
def choose_model(task_type: str, messages, system: str, tools=None) -> str:
    if task_type in {"classify", "extract", "tag", "route"} and not tools:
        # count_tokens is free and rate-limited separately from message creation.
        # Recount per target model: token counts are NOT comparable across
        # Claude tokenizer generations.
        count = client.messages.count_tokens(
            model=CHEAP, system=system, messages=messages
        ).input_tokens
        if count < 150_000:  # leave headroom inside Haiku's 200K window
            return CHEAP
        return MID
    if task_type in {"reason", "plan", "generate_code"}:
        return TOP
    return MID
Do not reuse a token count measured on one model to make a decision about another. Anthropic's documentation notes that Claude 4.7 and later models use a newer tokenizer that produces roughly 30% more tokens for the same text, and advises recounting against the model you plan to use. A routing rule calibrated on one tokenizer generation will misjudge context fit on another.

Pattern 2: cascade — cheap first, escalate on failure

A cascade sends everything to the cheap model, checks the result, and reruns on an expensive model when the check fails. It is more accurate than static routing because the decision is made after seeing an attempt rather than before. It is also the pattern people most often get wrong economically, because escalated requests pay twice.

The break-even maths

Let C_cheap be the cost of the cheap attempt, C_exp the cost of the expensive one, and r the fraction of requests that escalate. Always-expensive costs C_exp per request. The cascade costs C_cheap + r x C_exp. The cascade wins when:

C_cheap + (r x C_exp)  <  C_exp
 
           =>   r  <  1 - (C_cheap / C_exp)
 
Haiku 4.5 -> Opus 4.8, same token counts, cost ratio 1:5:
 
           r  <  1 - 0.2  =  0.80
 
Sonnet 5 -> Opus 4.8, cost ratio 2:5:
 
           r  <  1 - 0.4  =  0.60

An 80% escalation ceiling sounds like enormous headroom, and on tokens alone it is. But the money break-even is not the decision break-even, because escalation costs more than money:

  • **Latency is additive on escalation.** An escalated request waits for the cheap model to fail before the expensive one starts. At a 30% escalation rate, roughly a third of your users see something close to double the latency. That is a p90 regression that will not show up in your average.
  • **Your prompt cache fragments.** Each model keeps its own cache. Splitting traffic 70/30 means each model's prefix is warmed less often, and with a five-minute TTL a low-traffic tier can drop below its hit threshold entirely. In a lopsided split, the minority model may write its cache on nearly every request at a 1.25x multiplier and read it almost never.
  • **The minimum cacheable prefix differs by tier.** Claude Haiku 4.5 needs 4,096 tokens to cache; Opus 4.8 and Sonnet 5 need 1,024. A shared 2,000-token system prompt caches on the higher tiers and silently does not cache on Haiku.
  • **Tool overhead is not uniform.** The tool-use system prompt that the API injects when tools are present costs a different number of tokens on each model. On the figures published in August 2026, it is smaller on Opus 4.8 than on Haiku 4.5 — so a tool-heavy request narrows the tier gap slightly rather than widening it.
A practical planning rule: treat escalation rates above about 30% as a signal that the cheap tier is the wrong model for this task, regardless of what the token arithmetic says. At that point you are paying two latencies and running two eval suites to save a fraction you could get more cleanly by lowering effort on a single model.

Confidence signals that actually work

The cascade lives or dies on the escalation check. The single worst check is asking the model how confident it is in prose — self-reported confidence from a small model is poorly calibrated and tends to be uniformly high. Prefer signals that are mechanical and verifiable outside the model.

Signal How you get it Quality
Downstream validation failure Does the SQL parse? Does the code compile? Does the ID exist? Best — objective and free
Schema violation Structured output rejected, or a required field missing Excellent
Truncation `stop_reason == "max_tokens"` Excellent — a clear signal the task overflowed
Explicit abstention A schema field the model can set to decline, e.g. `insufficient_information` Good — you are asking for a decision, not a feeling
Enum-constrained confidence A `confidence` field restricted to `high`/`medium`/`low` in a strict schema Fair — calibrate it against ground truth before trusting it
Retrieval score Similarity of the supporting context you passed in Good, and available before you call the model at all
"Are you sure?" in prose Free-text self-assessment Poor — do not build on it
import anthropic
from pydantic import BaseModel
from typing import Literal
 
client = anthropic.Anthropic()
 
class Triage(BaseModel):
    category: Literal["billing", "technical", "account", "other"]
    # An explicit abstention channel. The model is deciding, not introspecting.
    sufficient_information: bool
 
def cascaded_triage(ticket: str) -> tuple[Triage, str]:
    cheap = client.messages.parse(
        model="claude-haiku-4-5",
        max_tokens=256,
        messages=[{"role": "user", "content": ticket}],
        output_format=Triage,
    )
    result = cheap.parsed_output
 
    # Escalate on abstention or on truncation — both are mechanical checks.
    if result.sufficient_information and cheap.stop_reason != "max_tokens":
        return result, "claude-haiku-4-5"
 
    strong = client.messages.parse(
        model="claude-opus-5",
        max_tokens=1024,
        messages=[{"role": "user", "content": ticket}],
        output_format=Triage,
    )
    return strong.parsed_output, "claude-opus-5"

Return the model that served the request, and log it. Escalation rate is the single most important operational metric a cascade has — it is both your cost driver and your earliest warning that input distribution has shifted.

Pattern 3: an LLM router

You can also ask a cheap model to classify the request and pick a tier. This is the most flexible pattern and the one that most often fails to pay for itself, because the router call sits on the critical path of every single request — including the ones that were always going to the cheap tier anyway.

It is worth it only when three things hold at once: the router is dramatically cheaper than the average request it routes (a 200-token classification against a 20,000-token task, not against a 500-token one), the routing decision genuinely cannot be made from static signals, and the added latency is acceptable on every request rather than just escalated ones. If you are considering an LLM router, first check whether a static rule on task type and input size gets you most of the way there. It usually does.

Implementing this without writing a router

Everything above is routing policy. The plumbing — a unified interface across providers, fallback chains, retries, per-model rate limits, per-key budgets and cost accounting — is a solved problem, and writing it yourself is rarely a good use of time. The AI Workshack guide to LiteLLM covers standing up the proxy with fallbacks and cost tracking, and the follow-up on running the LiteLLM proxy in production covers load balancing, caching and team access. Put your policy in that layer rather than scattering model names through application code.

The practical benefit is that model choice becomes configuration. When a new tier ships, you change a routing rule and re-run your evals rather than grepping for model IDs across a codebase.

Measuring the quality loss (this is the actual work)

Routing without measurement is not cost optimisation, it is a silent quality cut with a cost saving attached. The saving is visible on the invoice within days; the regression shows up in churn several months later. Build the measurement first.

  1. **Freeze an eval set from real traffic.** A few hundred real inputs, sampled to reflect your actual distribution — not a hand-picked set of easy cases. Include the awkward ones you would rather not look at.
  2. **Establish a reference.** Run the expensive model over the set and treat its outputs as the baseline. You are measuring the delta from what you ship today, not from an abstract notion of correctness.
  3. **Score agreement, not vibes.** For classification, exact match. For extraction, field-level accuracy. For open text, a judge model with a rubric, validated against a human-labelled subset before you trust it. Write down the metric before you run the experiment.
  4. **Break results down by segment.** Aggregate accuracy hides everything that matters. A cheap model that scores 94% overall may be at 99% on your biggest customer type and 61% on a smaller one. The average will not tell you, and the smaller segment is often the one that complains loudest.
  5. **Shadow-run before switching.** Send production traffic to both models, serve the expensive result, log both. You get real-distribution comparison data with zero user risk. It costs you double for the duration — budget for it as an experiment.
  6. **Set a quality floor and treat it as a release gate.** Decide in advance what agreement rate is acceptable per segment, and hold the routing change to it exactly as you would hold a code change to a test suite.

For the trace-and-score infrastructure — capturing model, prompt version, token counts and eval scores per request so you can slice them later — the AI Workshack guides to LangFuse and to the LangSmith comparison cover the setup. The key requirement is that model identity is a first-class dimension in your traces, so "which model served this" is a filter rather than an archaeology project.

Watch for these after you ship

  • **Escalation rate drift.** A cascade that escalated 20% of requests in testing and 55% in production means your eval set did not match reality. It is also costing you more than the single-model baseline.
  • **Cache hit rate falling on both tiers.** The classic routing side effect. Check it explicitly after every routing change.
  • **Prompt divergence.** Two models, two prompts, and over a few months only one of them gets maintained. Version the prompts together, and re-run both eval suites on every prompt change.
  • **Quality complaints that do not map to errors.** The cheap tier rarely fails loudly. It fails by being slightly worse — a little less specific, a little more generic. That shows up in user sentiment long before it shows up in a metric, so keep a qualitative review loop.

The honest summary

Approach Typical saving What it costs you
Lower `effort` on one model Meaningful, workload-dependent Some reasoning depth. One parameter, one eval suite, cache intact.
Static routing on task type Up to the tier ratio on routed traffic Two prompts, two eval suites, split cache.
Cascade with escalation Good below ~30% escalation Latency on escalated requests, double cost on those requests, most complexity.
LLM router Situational A model call on every request's critical path. Usually not worth it.

Start at the top of that table and work down. Effort tuning is nearly free and reversible. Static routing is a day of work. A cascade is a system with its own failure modes and its own operational metric. Each step down buys a bit more saving for a lot more complexity, and most teams should stop before the bottom row.

Model IDs, tier prices, context windows and parameter support in this article were verified against Anthropic documentation in August 2026. New tiers ship regularly and pricing moves — re-check the current lineup and re-run your evals before rolling a routing change forward.