Where your tokens go, how to see them, and how to spend fewer of them — prompt caching, batching, model tiering, context engineering, agent cost control, and usage monitoring. Provider-neutral, with the tradeoffs named.
What a token actually is, what you are billed for, and how to count before you pay.
Four counters, four different prices, and a set of multipliers that stack. Here is how to turn a usage object into a cost number you can actually track per request.
The rule of thumb everyone quotes is wrong in the direction that costs you money. Here is what your provider is really metering, and how to measure it before the invoice arrives.
The built-in discounts: prompt caching, batch APIs, model tiering, and output control.
Anthropic Message Batches and the OpenAI Batch API — the discount, the real SLA, what is safely deferrable, and the failure modes that eat the saving.
Why output costs 5x input, what max_tokens actually does (it is not a budget), which concision instructions work, and how to measure the difference.
Cascade patterns, routing heuristics that do not need an LLM, the break-even maths on escalation, and how to prove you did not quietly degrade quality.
Three providers, three completely different caching contracts — and the cases where turning caching on makes your bill go up.
Shrinking what you send — context windows, prompt and tool-schema bloat, and RAG economics.
A long conversation is not expensive because it is long. It is expensive because you re-send all of it, every single turn.
In most RAG systems, retrieved context is over 80% of your input tokens and none of it caches. Here is how to retrieve less, pay less, and often answer better.
Your history grows and gets trimmed. Your system prompt and tool schemas are re-sent in full, unchanged, on every single request — forever.
Why agent loops burn tokens, and how to keep coding agents and multi-agent systems affordable.
Every subagent call is a full LLM invocation with its own context. Whether that saves money or triples it comes down to one question: how much context do your subagents share?
Claude Code, Cursor, Copilot and Replit all meter the same thing — the size of the context you carry. Here are the habits that shrink it, and what each one costs you.
Most builders learn this from an invoice. An agent re-sends its entire conversation on every single turn, so cost grows with the square of the turn count, not in a straight line.
Seeing where the money goes: usage APIs, per-user attribution, budgets and alerts.
Hard caps versus soft alerts, per-tenant budgets, rate limits as a cost control, detecting agent runaway loops, and the escalation runbook that turns a spike into a ticket instead of a board conversat...
Tagging strategy, computing cost from usage fields, unit economics, and chargeback that survives an audit — including the awkward cases of cached and batched calls.
Native usage APIs, observability platforms, gateway metering and OpenTelemetry — what each layer actually gives you, what it costs you, and how to pick without locking yourself in.
New guides drop regularly. Get them in your inbox — no noise, just signal.