Shrinking what you send — context windows, prompt and tool-schema bloat, and RAG economics.
A long conversation is not expensive because it is long. It is expensive because you re-send all of it, every single turn.
In most RAG systems, retrieved context is over 80% of your input tokens and none of it caches. Here is how to retrieve less, pay less, and often answer better.
Your history grows and gets trimmed. Your system prompt and tool schemas are re-sent in full, unchanged, on every single request — forever.
New guides drop regularly. Get them in your inbox — no noise, just signal.