Hard caps versus soft alerts, per-tenant budgets, rate limits as a cost control, detecting agent runaway loops, and the escalation runbook that turns a spike into a ticket instead of a board conversation.
Most teams discover a runaway LLM bill in one of two ways: a monthly invoice, or a provider-side rate limit that takes production down. Both are late. The gap between the first anomalous request and the moment a human notices is typically measured in days, and in that window an agent loop or a bad deploy can spend more than the feature will earn in a year.
Closing that gap is not one control. It is a layered system of caps, limits, detectors and a runbook, and each layer buys you something specific at a specific cost. This article works through all four.
All prices in this article are as of August 2026 and are quoted only to make the arithmetic concrete. Verify current pricing against the provider's own pricing page before you put a number in a spreadsheet, a budget, or a customer contract.The control hierarchy
| Control | Enforcement point | Latency to effect | Blast radius | The tradeoff |
|---|---|---|---|---|
| Hard cap | Gateway or app middleware | Immediate | Everything under the cap stops | It is an availability incident by design |
| Soft alert | Monitoring system | Minutes to hours | None | Only works if a human is watching |
| Rate limit | Gateway or provider | Immediate | Bounded to the limited scope | Bounds the rate, not the total |
| Circuit breaker | Application | Immediate | One feature or one tenant | Needs a defined degraded mode to be safe |
| Provider spend limit | Provider console | Immediate | Entire organisation | Coarsest possible control; use only as a backstop |
The mistake almost everyone makes is choosing between hard caps and soft alerts. You need both, at different scopes: hard caps where the blast radius is acceptable, alerts everywhere else.
Hard caps: only where you have a degraded mode
A hard cap on a customer-facing path without a fallback is a self-inflicted outage with a nicer name. Before enabling one, define what happens at the boundary:
- Downgrade the model. Route from claude-opus-5 to claude-sonnet-5, or from claude-sonnet-5 to claude-haiku-4-5. Cost drops materially; quality drops somewhat. Users keep working.
- Disable optional enrichment. Skip the summary, the suggested replies, the auto-tagging. The core action still completes.
- Queue for batch. Move the work to the batch tier, which as of August 2026 is documented at a 50% discount on both input and output. Latency goes from seconds to hours - fine for a nightly report, not for a chat box.
- Refuse, with an honest message. Legitimate for trial tenants and internal tools. Never for a paying customer without a contractual basis.
Write the degraded mode first and the cap second. A cap whose only behaviour is a 429 will be disabled by the first person paged at 3am, and it will stay disabled.Where budgets can actually live
Provider console
Organisation-wide only. Useful as a final backstop against catastrophic runaway, useless for anything granular, and it fails closed for every workload at once. Set it well above expected spend and treat it as a fuse, not a control.
Gateway
The natural home for per-key and per-tenant budgets, because the gateway sees every call and can refuse one. If you already run LiteLLM, its virtual keys carry max_budget in USD with a budget_duration reset window, and team-level budgets and soft-budget alert thresholds work the same way — our existing LiteLLM proxy article covers the configuration in detail.
Portkey offers the same category of control from the managed side: budget limits on API keys, workspaces and providers, expressible in either cost or tokens, with an optional alert threshold that fires while the key still works, and periodic reset options including weekly (Sundays at 00:00 UTC) and monthly (the 1st at 00:00 UTC). Note the commercial caveat before you design around it: as of August 2026 budget limits are documented as available to Enterprise and selected Pro customers rather than to all paid plans. Verify your plan covers it.
Application middleware
The most flexible and the most code. This is where you express things a gateway cannot: budget per feature per tenant, different caps for trial versus paid, a per-session ceiling on an agent run.
The hard part is that you cannot know a call's cost before making it. The workable pattern is reserve-then-settle: reserve a worst-case estimate before the call, settle to the actual afterwards.
from decimal import Decimal
class BudgetExceeded(Exception):
pass
def worst_case_usd(model: str, input_tokens: int, max_tokens: int) -> Decimal:
r = PRICES[model]
return (Decimal(input_tokens) * Decimal(r["in"])
+ Decimal(max_tokens) * Decimal(r["out"])) / Decimal(1_000_000)
def guarded_call(prompt: str, model: str, max_tokens: int):
ctx = current_cost_context()
est_in = count_tokens(model, prompt) # provider token-count endpoint
reservation = worst_case_usd(model, est_in, max_tokens)
# Atomic reserve against the tenant's remaining window budget.
if not budget_store.reserve(ctx.tenant_id, reservation):
raise BudgetExceeded(ctx.tenant_id)
try:
resp = client.messages.create(
model=model, max_tokens=max_tokens,
messages=[{"role": "user", "content": prompt}],
metadata={"user_id": ctx.user_hash},
)
except Exception:
budget_store.release(ctx.tenant_id, reservation)
raise
actual = Decimal(str(price_request(model, resp.usage)["cost_usd"]))
budget_store.settle(ctx.tenant_id, reservation, actual)
return resp
Two things make this practical. First, max_tokens is what bounds the reservation — without it the worst case is the full context window and every reservation is absurd. With max_tokens=1024 on claude-sonnet-5 at $10 per million output tokens (as of August 2026), the output side of a single call cannot exceed roughly $0.0102, which makes the reservation tight enough to be useful. Second, reservations must expire, or a crashed process leaks budget until someone notices.
The tradeoff: over-reservation rejects legitimate traffic near the boundary. Tune by measuring the ratio of settled to reserved. If it sits below about 0.3, your max_tokens values are too loose and the budget is being enforced against a fiction.
Rate limits as a cost control
Rate limits are the most under-used cost control because most teams inherit them as an availability mechanism and never revisit the unit.
Requests per minute is a poor cost proxy. One request can be a 200-token classification or a 900,000-token document analysis, and they differ in price by four orders of magnitude. Limit on tokens per minute where you can, and on spend where the tool supports it.
Helicone expresses this directly with a Helicone-RateLimit-Policy header in the form [quota];w=[time window];u=[unit];s=[segment]. The window is in seconds with a minimum of 60. Setting u=cents switches from counting requests to counting spend, and s=user segments per Helicone-User-Id, while s=[custom property] segments on any custom property such as an organisation.
# 1,000 requests per hour, globally
Helicone-RateLimit-Policy: 1000;w=3600
# $5.00 per hour, per end user
Helicone-RateLimit-Policy: 500;w=3600;u=cents;s=user
# $50.00 per day, per tenant (custom property named organization)
Helicone-RateLimit-Policy: 5000;w=86400;u=cents;s=organization
The tradeoff with cost-based limits is that they are enforced on observed spend, which means they are inherently reactive: a single very large request can carry a user past the limit in one call rather than being blocked before it. Pair a cost limit with a per-request max_tokens ceiling so the overshoot is bounded.
| Limit on | Catches | Misses |
|---|---|---|
| Requests per minute | Loops, retry storms, scrapers | Any expensive single call |
| Tokens per minute | Context bloat, large documents | Model-mix changes (Haiku and Opus tokens cost differently) |
| Spend per window | Everything, in the right unit | Enforcement lags by one request |
| Concurrency | Fan-out and parallel agent explosions | Slow sustained burn |
Detecting agent runaway loops
Agentic workloads are where cost incidents actually come from now. The failure mode is specific: an agent that cannot make progress does not stop, it retries — and because each turn re-sends the accumulated context, cost per turn rises as the loop continues. The spend curve is superlinear, which is exactly why a daily check is too slow.
The measurable signals, in rough order of how early they fire:
- Steps per session above the p99 for that agent. The cleanest signal and usually the first to move.
- Repeated identical tool calls. The same tool with the same arguments three times in one session is a loop, not a strategy.
- Context growth per turn that is not decelerating. A healthy agent's context growth flattens as it converges; a stuck one grows linearly forever.
- Tool-call-to-inference ratio collapsing toward zero - the agent is thinking without acting.
- Cumulative session cost crossing a multiple of the median session for that workflow.
The OpenTelemetry GenAI conventions define counters for exactly this shape of workload: gen_ai.invoke_agent.inference_calls and gen_ai.invoke_agent.tool_calls, alongside gen_ai.invoke_agent.duration and gen_ai.execute_tool.duration. If you are already emitting those, the detection is a query rather than a project. Note that these conventions are marked Development as of August 2026 and the names may still change.
Detection is necessary but not sufficient. Every agent run needs hard stops enforced in the loop itself, because by the time an alert fires the money is spent.
from dataclasses import dataclass, field
from decimal import Decimal
import time
@dataclass
class SessionGuard:
max_steps: int = 25
max_cost_usd: Decimal = Decimal("2.00")
max_wall_seconds: float = 300.0
max_identical_calls: int = 3
steps: int = 0
spent: Decimal = Decimal("0")
started: float = field(default_factory=time.monotonic)
_calls: dict[tuple, int] = field(default_factory=dict)
def check(self) -> None:
if self.steps >= self.max_steps:
raise RunawayDetected("step_limit", self.steps)
if self.spent >= self.max_cost_usd:
raise RunawayDetected("cost_limit", float(self.spent))
if time.monotonic() - self.started >= self.max_wall_seconds:
raise RunawayDetected("wall_clock", self.max_wall_seconds)
def record_step(self, cost_usd: Decimal, tool: str, args_hash: str) -> None:
self.steps += 1
self.spent += cost_usd
key = (tool, args_hash)
self._calls[key] = self._calls.get(key, 0) + 1
if self._calls[key] > self.max_identical_calls:
raise RunawayDetected("repeated_tool_call", key)
self.check()
The tradeoff is genuine: a step limit that is too tight truncates legitimate hard tasks, and users experience that as the agent giving up. Set the limits from your own distribution — p99 of successful sessions, not a round number — and emit a metric every time a guard fires so you can tell truncation from protection.
What to alert on
Most teams alert on total daily spend, which is the single least useful signal available: it fires a day late and tells you nothing about cause. These are the signals that actually lead somewhere.
| Signal | Why it matters | Detection | False-positive risk |
|---|---|---|---|
| Spend velocity ($/hour vs trailing baseline) | The earliest general-purpose signal | EWMA of hourly spend, alert on ratio to a 7-day same-hour baseline | High on launches and marketing pushes |
| Input tokens per request, p50 and p95, by feature | Catches prompt or context bloat from a deploy | Compare against the previous release | Low - very specific |
| Cache hit rate (cache read / total input tokens) | A cache that stops working can multiply input cost | Alert on a relative drop, e.g. below 70% of the 7-day median | Medium - moves with traffic mix |
| Attempts per logical request | Retry storms bill in full on every attempt | Alert above ~1.2 | Low |
| Failed-but-billed spend | Paid input on responses you discarded | Cost where status is error, as a share of total | Low |
| Steps or cost per agent session, p99 | The runaway loop signal | Per-workflow percentile tracking | Medium on genuinely hard tasks |
| Spend concentration (top tenant share) | One customer becoming the whole bill | Top-1 and top-5 share of daily spend | Low |
| New model or new key appearing | Unreviewed spend paths | Diff the set of distinct models and key IDs daily | Very low, and worth knowing anyway |
Anomaly detection without an ML project
You do not need a model. Three techniques cover almost everything:
- Seasonal baselines. Compare each hour to the same hour on the same weekday over the past few weeks. LLM traffic is intensely diurnal and weekly; a flat threshold will either page you every Monday morning or never.
- Robust z-scores. Use median and median absolute deviation rather than mean and standard deviation. A single $400 hour drags a mean so far that the next incident looks normal.
- Ratio metrics, not absolutes. Cost per request, tokens per request, cache hit rate. These are stable across traffic growth, which means the threshold you set in March still works in September.
The tradeoff with baseline-relative alerting is launch day: a legitimate 5x traffic increase looks exactly like an incident. Combine a relative condition with an absolute floor so small absolute numbers never page, and add a documented suppression window that a human opens deliberately before a launch rather than one that fires automatically.
Alert on cost you compute from usage fields, not on the provider's cost API. Provider cost data typically lags by minutes at best and, for cost endpoints specifically, is often daily-granularity only. It is the right source for reconciliation and the wrong source for detection.The escalation and runbook pattern
Alerts without a runbook produce a Slack channel that everyone mutes. Define severity by rate of spend and by who can stop it.
| Severity | Trigger (illustrative) | First response | Owner | Time-box |
|---|---|---|---|---|
| SEV-3 informational | Daily spend 1.5x baseline | Note in the weekly cost review | Cost owner | Next review |
| SEV-2 investigate | Hourly spend 3x baseline for 2 consecutive hours | Page during business hours; identify the top dimension | On-call engineer | 4 hours |
| SEV-1 contain | Hourly spend 10x baseline, or projected month-end above the approved budget | Page immediately; contain before diagnosing | On-call plus engineering manager | 30 minutes |
The runbook itself should be short enough to follow at 3am:
- Confirm it is real. Check that the spike is in your metering and not an instrumentation change - a new integration double-logging looks identical to a cost spike on the dashboard.
- Scope it. Break down by tenant, feature, model and environment, in that order. One of those four dimensions almost always explains it in a single query. If none of them do, the cause is probably in the unattributed bucket, which is its own finding.
- Contain before you understand. Apply the narrowest available control: disable the one feature flag, throttle the one tenant, revoke the one key. Containment does not require a diagnosis.
- Attribute. Once contained, produce the cost figure and the dimension. Bring a number to the incident channel, not an adjective.
- Remediate. Fix the cause. Common causes in practice: a prompt change that broke a cache breakpoint, a retry policy without a ceiling, an agent loop with no step limit, a CI job pointed at production credentials, a customer integration in a tight loop.
- Reconcile. When the provider cost data catches up, compare it to your computed figure for the incident window. A gap here is a second bug.
- Post-incident: add the detector. Every cost incident should end with a new signal or a tightened threshold, and with the guard limits updated if an agent was involved.
Number four is the one that changes outcomes. 'Spend is high' starts a debate. '$4,120 in six hours, 94% from tenant 8812's document pipeline, caused by a retry loop on a malformed PDF' ends one.Testing that any of this works
Alerting that has never fired is not alerting, it is decoration. Test it deliberately, on a schedule:
- Provision a burner key with a small budget and drive synthetic spend through it until the cap trips. Confirm the alert fires, reaches a human, and that the degraded mode behaved as designed.
- Break a cache breakpoint in staging and confirm the cache-hit-rate alert fires before the spend alert does. If it does not, your leading indicator is not leading.
- Run an agent against a task engineered to be unsolvable, and confirm the session guard stops it and emits the metric.
- Once a quarter, replay a past incident's traffic shape against the current thresholds. Traffic grows; thresholds set eighteen months ago usually no longer fire.
The weekly ritual that beats every dashboard
Thirty minutes, one owner, four numbers: total spend versus budget, cost per active user, the top three features by spend and their trend, and the size of the unattributed bucket. Then one question — what changed since last week — answered from the release log rather than from memory.
Automation catches the spikes. The weekly review catches the drift, which is where most of the money actually goes: a prompt that grew 400 tokens over a quarter, a default that quietly moved from claude-haiku-4-5 to claude-sonnet-5, a retry policy that was never tuned. None of those trip an anomaly detector. All of them show up in a ratio metric looked at by a human every week.
Budgets stop catastrophes. Alerts stop incidents. The review stops the slow bleed, and over a year it is usually the largest of the three.