Hard caps versus soft alerts, per-tenant budgets, rate limits as a cost control, detecting agent runaway loops, and the escalation runbook that turns a spike into a ticket instead of a board conversation.

Most teams discover a runaway LLM bill in one of two ways: a monthly invoice, or a provider-side rate limit that takes production down. Both are late. The gap between the first anomalous request and the moment a human notices is typically measured in days, and in that window an agent loop or a bad deploy can spend more than the feature will earn in a year.

Closing that gap is not one control. It is a layered system of caps, limits, detectors and a runbook, and each layer buys you something specific at a specific cost. This article works through all four.

All prices in this article are as of August 2026 and are quoted only to make the arithmetic concrete. Verify current pricing against the provider's own pricing page before you put a number in a spreadsheet, a budget, or a customer contract.

The control hierarchy

Control Enforcement point Latency to effect Blast radius The tradeoff
Hard cap Gateway or app middleware Immediate Everything under the cap stops It is an availability incident by design
Soft alert Monitoring system Minutes to hours None Only works if a human is watching
Rate limit Gateway or provider Immediate Bounded to the limited scope Bounds the rate, not the total
Circuit breaker Application Immediate One feature or one tenant Needs a defined degraded mode to be safe
Provider spend limit Provider console Immediate Entire organisation Coarsest possible control; use only as a backstop

The mistake almost everyone makes is choosing between hard caps and soft alerts. You need both, at different scopes: hard caps where the blast radius is acceptable, alerts everywhere else.

Hard caps: only where you have a degraded mode

A hard cap on a customer-facing path without a fallback is a self-inflicted outage with a nicer name. Before enabling one, define what happens at the boundary:

  • Downgrade the model. Route from claude-opus-5 to claude-sonnet-5, or from claude-sonnet-5 to claude-haiku-4-5. Cost drops materially; quality drops somewhat. Users keep working.
  • Disable optional enrichment. Skip the summary, the suggested replies, the auto-tagging. The core action still completes.
  • Queue for batch. Move the work to the batch tier, which as of August 2026 is documented at a 50% discount on both input and output. Latency goes from seconds to hours - fine for a nightly report, not for a chat box.
  • Refuse, with an honest message. Legitimate for trial tenants and internal tools. Never for a paying customer without a contractual basis.
Write the degraded mode first and the cap second. A cap whose only behaviour is a 429 will be disabled by the first person paged at 3am, and it will stay disabled.

Where budgets can actually live

Provider console

Organisation-wide only. Useful as a final backstop against catastrophic runaway, useless for anything granular, and it fails closed for every workload at once. Set it well above expected spend and treat it as a fuse, not a control.

Gateway

The natural home for per-key and per-tenant budgets, because the gateway sees every call and can refuse one. If you already run LiteLLM, its virtual keys carry max_budget in USD with a budget_duration reset window, and team-level budgets and soft-budget alert thresholds work the same way — our existing LiteLLM proxy article covers the configuration in detail.

Portkey offers the same category of control from the managed side: budget limits on API keys, workspaces and providers, expressible in either cost or tokens, with an optional alert threshold that fires while the key still works, and periodic reset options including weekly (Sundays at 00:00 UTC) and monthly (the 1st at 00:00 UTC). Note the commercial caveat before you design around it: as of August 2026 budget limits are documented as available to Enterprise and selected Pro customers rather than to all paid plans. Verify your plan covers it.

Application middleware

The most flexible and the most code. This is where you express things a gateway cannot: budget per feature per tenant, different caps for trial versus paid, a per-session ceiling on an agent run.

The hard part is that you cannot know a call's cost before making it. The workable pattern is reserve-then-settle: reserve a worst-case estimate before the call, settle to the actual afterwards.

from decimal import Decimal
 
class BudgetExceeded(Exception):
    pass
 
def worst_case_usd(model: str, input_tokens: int, max_tokens: int) -> Decimal:
    r = PRICES[model]
    return (Decimal(input_tokens) * Decimal(r["in"])
            + Decimal(max_tokens) * Decimal(r["out"])) / Decimal(1_000_000)
 
def guarded_call(prompt: str, model: str, max_tokens: int):
    ctx = current_cost_context()
    est_in = count_tokens(model, prompt)          # provider token-count endpoint
    reservation = worst_case_usd(model, est_in, max_tokens)
 
    # Atomic reserve against the tenant's remaining window budget.
    if not budget_store.reserve(ctx.tenant_id, reservation):
        raise BudgetExceeded(ctx.tenant_id)
    try:
        resp = client.messages.create(
            model=model, max_tokens=max_tokens,
            messages=[{"role": "user", "content": prompt}],
            metadata={"user_id": ctx.user_hash},
        )
    except Exception:
        budget_store.release(ctx.tenant_id, reservation)
        raise
    actual = Decimal(str(price_request(model, resp.usage)["cost_usd"]))
    budget_store.settle(ctx.tenant_id, reservation, actual)
    return resp
 

Two things make this practical. First, max_tokens is what bounds the reservation — without it the worst case is the full context window and every reservation is absurd. With max_tokens=1024 on claude-sonnet-5 at $10 per million output tokens (as of August 2026), the output side of a single call cannot exceed roughly $0.0102, which makes the reservation tight enough to be useful. Second, reservations must expire, or a crashed process leaks budget until someone notices.

The tradeoff: over-reservation rejects legitimate traffic near the boundary. Tune by measuring the ratio of settled to reserved. If it sits below about 0.3, your max_tokens values are too loose and the budget is being enforced against a fiction.

Rate limits as a cost control

Rate limits are the most under-used cost control because most teams inherit them as an availability mechanism and never revisit the unit.

Requests per minute is a poor cost proxy. One request can be a 200-token classification or a 900,000-token document analysis, and they differ in price by four orders of magnitude. Limit on tokens per minute where you can, and on spend where the tool supports it.

Helicone expresses this directly with a Helicone-RateLimit-Policy header in the form [quota];w=[time window];u=[unit];s=[segment]. The window is in seconds with a minimum of 60. Setting u=cents switches from counting requests to counting spend, and s=user segments per Helicone-User-Id, while s=[custom property] segments on any custom property such as an organisation.

# 1,000 requests per hour, globally
Helicone-RateLimit-Policy: 1000;w=3600
 
# $5.00 per hour, per end user
Helicone-RateLimit-Policy: 500;w=3600;u=cents;s=user
 
# $50.00 per day, per tenant (custom property named organization)
Helicone-RateLimit-Policy: 5000;w=86400;u=cents;s=organization
 

The tradeoff with cost-based limits is that they are enforced on observed spend, which means they are inherently reactive: a single very large request can carry a user past the limit in one call rather than being blocked before it. Pair a cost limit with a per-request max_tokens ceiling so the overshoot is bounded.

Limit on Catches Misses
Requests per minute Loops, retry storms, scrapers Any expensive single call
Tokens per minute Context bloat, large documents Model-mix changes (Haiku and Opus tokens cost differently)
Spend per window Everything, in the right unit Enforcement lags by one request
Concurrency Fan-out and parallel agent explosions Slow sustained burn

Detecting agent runaway loops

Agentic workloads are where cost incidents actually come from now. The failure mode is specific: an agent that cannot make progress does not stop, it retries — and because each turn re-sends the accumulated context, cost per turn rises as the loop continues. The spend curve is superlinear, which is exactly why a daily check is too slow.

The measurable signals, in rough order of how early they fire:

  • Steps per session above the p99 for that agent. The cleanest signal and usually the first to move.
  • Repeated identical tool calls. The same tool with the same arguments three times in one session is a loop, not a strategy.
  • Context growth per turn that is not decelerating. A healthy agent's context growth flattens as it converges; a stuck one grows linearly forever.
  • Tool-call-to-inference ratio collapsing toward zero - the agent is thinking without acting.
  • Cumulative session cost crossing a multiple of the median session for that workflow.

The OpenTelemetry GenAI conventions define counters for exactly this shape of workload: gen_ai.invoke_agent.inference_calls and gen_ai.invoke_agent.tool_calls, alongside gen_ai.invoke_agent.duration and gen_ai.execute_tool.duration. If you are already emitting those, the detection is a query rather than a project. Note that these conventions are marked Development as of August 2026 and the names may still change.

Detection is necessary but not sufficient. Every agent run needs hard stops enforced in the loop itself, because by the time an alert fires the money is spent.

from dataclasses import dataclass, field
from decimal import Decimal
import time
 
@dataclass
class SessionGuard:
    max_steps: int = 25
    max_cost_usd: Decimal = Decimal("2.00")
    max_wall_seconds: float = 300.0
    max_identical_calls: int = 3
 
    steps: int = 0
    spent: Decimal = Decimal("0")
    started: float = field(default_factory=time.monotonic)
    _calls: dict[tuple, int] = field(default_factory=dict)
 
    def check(self) -> None:
        if self.steps >= self.max_steps:
            raise RunawayDetected("step_limit", self.steps)
        if self.spent >= self.max_cost_usd:
            raise RunawayDetected("cost_limit", float(self.spent))
        if time.monotonic() - self.started >= self.max_wall_seconds:
            raise RunawayDetected("wall_clock", self.max_wall_seconds)
 
    def record_step(self, cost_usd: Decimal, tool: str, args_hash: str) -> None:
        self.steps += 1
        self.spent += cost_usd
        key = (tool, args_hash)
        self._calls[key] = self._calls.get(key, 0) + 1
        if self._calls[key] > self.max_identical_calls:
            raise RunawayDetected("repeated_tool_call", key)
        self.check()
 

The tradeoff is genuine: a step limit that is too tight truncates legitimate hard tasks, and users experience that as the agent giving up. Set the limits from your own distribution — p99 of successful sessions, not a round number — and emit a metric every time a guard fires so you can tell truncation from protection.

What to alert on

Most teams alert on total daily spend, which is the single least useful signal available: it fires a day late and tells you nothing about cause. These are the signals that actually lead somewhere.

Signal Why it matters Detection False-positive risk
Spend velocity ($/hour vs trailing baseline) The earliest general-purpose signal EWMA of hourly spend, alert on ratio to a 7-day same-hour baseline High on launches and marketing pushes
Input tokens per request, p50 and p95, by feature Catches prompt or context bloat from a deploy Compare against the previous release Low - very specific
Cache hit rate (cache read / total input tokens) A cache that stops working can multiply input cost Alert on a relative drop, e.g. below 70% of the 7-day median Medium - moves with traffic mix
Attempts per logical request Retry storms bill in full on every attempt Alert above ~1.2 Low
Failed-but-billed spend Paid input on responses you discarded Cost where status is error, as a share of total Low
Steps or cost per agent session, p99 The runaway loop signal Per-workflow percentile tracking Medium on genuinely hard tasks
Spend concentration (top tenant share) One customer becoming the whole bill Top-1 and top-5 share of daily spend Low
New model or new key appearing Unreviewed spend paths Diff the set of distinct models and key IDs daily Very low, and worth knowing anyway

Anomaly detection without an ML project

You do not need a model. Three techniques cover almost everything:

  1. Seasonal baselines. Compare each hour to the same hour on the same weekday over the past few weeks. LLM traffic is intensely diurnal and weekly; a flat threshold will either page you every Monday morning or never.
  2. Robust z-scores. Use median and median absolute deviation rather than mean and standard deviation. A single $400 hour drags a mean so far that the next incident looks normal.
  3. Ratio metrics, not absolutes. Cost per request, tokens per request, cache hit rate. These are stable across traffic growth, which means the threshold you set in March still works in September.

The tradeoff with baseline-relative alerting is launch day: a legitimate 5x traffic increase looks exactly like an incident. Combine a relative condition with an absolute floor so small absolute numbers never page, and add a documented suppression window that a human opens deliberately before a launch rather than one that fires automatically.

Alert on cost you compute from usage fields, not on the provider's cost API. Provider cost data typically lags by minutes at best and, for cost endpoints specifically, is often daily-granularity only. It is the right source for reconciliation and the wrong source for detection.

The escalation and runbook pattern

Alerts without a runbook produce a Slack channel that everyone mutes. Define severity by rate of spend and by who can stop it.

Severity Trigger (illustrative) First response Owner Time-box
SEV-3 informational Daily spend 1.5x baseline Note in the weekly cost review Cost owner Next review
SEV-2 investigate Hourly spend 3x baseline for 2 consecutive hours Page during business hours; identify the top dimension On-call engineer 4 hours
SEV-1 contain Hourly spend 10x baseline, or projected month-end above the approved budget Page immediately; contain before diagnosing On-call plus engineering manager 30 minutes

The runbook itself should be short enough to follow at 3am:

  1. Confirm it is real. Check that the spike is in your metering and not an instrumentation change - a new integration double-logging looks identical to a cost spike on the dashboard.
  2. Scope it. Break down by tenant, feature, model and environment, in that order. One of those four dimensions almost always explains it in a single query. If none of them do, the cause is probably in the unattributed bucket, which is its own finding.
  3. Contain before you understand. Apply the narrowest available control: disable the one feature flag, throttle the one tenant, revoke the one key. Containment does not require a diagnosis.
  4. Attribute. Once contained, produce the cost figure and the dimension. Bring a number to the incident channel, not an adjective.
  5. Remediate. Fix the cause. Common causes in practice: a prompt change that broke a cache breakpoint, a retry policy without a ceiling, an agent loop with no step limit, a CI job pointed at production credentials, a customer integration in a tight loop.
  6. Reconcile. When the provider cost data catches up, compare it to your computed figure for the incident window. A gap here is a second bug.
  7. Post-incident: add the detector. Every cost incident should end with a new signal or a tightened threshold, and with the guard limits updated if an agent was involved.
Number four is the one that changes outcomes. 'Spend is high' starts a debate. '$4,120 in six hours, 94% from tenant 8812's document pipeline, caused by a retry loop on a malformed PDF' ends one.

Testing that any of this works

Alerting that has never fired is not alerting, it is decoration. Test it deliberately, on a schedule:

  • Provision a burner key with a small budget and drive synthetic spend through it until the cap trips. Confirm the alert fires, reaches a human, and that the degraded mode behaved as designed.
  • Break a cache breakpoint in staging and confirm the cache-hit-rate alert fires before the spend alert does. If it does not, your leading indicator is not leading.
  • Run an agent against a task engineered to be unsolvable, and confirm the session guard stops it and emits the metric.
  • Once a quarter, replay a past incident's traffic shape against the current thresholds. Traffic grows; thresholds set eighteen months ago usually no longer fire.

The weekly ritual that beats every dashboard

Thirty minutes, one owner, four numbers: total spend versus budget, cost per active user, the top three features by spend and their trend, and the size of the unattributed bucket. Then one question — what changed since last week — answered from the release log rather than from memory.

Automation catches the spikes. The weekly review catches the drift, which is where most of the money actually goes: a prompt that grew 400 tokens over a quarter, a default that quietly moved from claude-haiku-4-5 to claude-sonnet-5, a retry policy that was never tuned. None of those trip an anomaly detector. All of them show up in a ratio metric looked at by a human every week.

Budgets stop catastrophes. Alerts stop incidents. The review stops the slow bleed, and over a year it is usually the largest of the three.