The built-in discounts: prompt caching, batch APIs, model tiering, and output control.
Anthropic Message Batches and the OpenAI Batch API — the discount, the real SLA, what is safely deferrable, and the failure modes that eat the saving.
Why output costs 5x input, what max_tokens actually does (it is not a budget), which concision instructions work, and how to measure the difference.
Cascade patterns, routing heuristics that do not need an LLM, the break-even maths on escalation, and how to prove you did not quietly degrade quality.
Three providers, three completely different caching contracts — and the cases where turning caching on makes your bill go up.
New guides drop regularly. Get them in your inbox — no noise, just signal.