Claude Code, Cursor, Copilot and Replit all meter the same thing — the size of the context you carry. Here are the habits that shrink it, and what each one costs you.

There are two ways to run out of an agentic coding tool. On a subscription plan you hit a usage limit mid-task and lose an afternoon. On an API or cloud-provider plan you get a bill that makes someone ask questions. Both come from the same place: the size of the context your session is carrying, multiplied by the number of times it gets sent.

What follows is a set of working habits, ranked roughly by impact, each with the thing it costs you — useful whether you never see a dollar figure or you own the invoice. Where a command or setting is named it has been checked against current vendor documentation; where a practice is real but the syntax varies by tool or version, it is described generically on purpose.

Figures, commands, and plan mechanics here are as of August 2026 and were verified against Anthropic, Cursor, and GitHub primary documentation. Agentic coding tools ship weekly — confirm command names and limits in your own installed version before relying on them.

First, understand what you are actually being metered on

Every one of these tools sends the whole conversation to the model on every turn, and sends another request each time the agent uses a tool. That is not a quirk of any one product; it is how a stateless chat API becomes an agent. The practical consequence, in Claude Code's own words, is that a one-line question in a session that has been open all day still draws usage for the whole conversation. The billing surfaces differ; the meter underneath is the same.

Tool How it is metered (as of Aug 2026) What that means for you
Claude Code API token consumption; on subscription plans a per-seat allowance on a rolling five-hour window plus a weekly window Context size drives both the bill and how fast you hit the window
Cursor Per token (input, cache read, cache write, output) drawn from two monthly pools: Cursor Models and Other Models Model choice changes how fast the pool drains, not just the rate
GitHub Copilot Token-based, denominated in GitHub AI Credits at 1 credit = $0.01, priced per million tokens by model Completions and next edit suggestions are unbilled and unlimited on paid plans; agent and chat work is metered
Replit Agent Credit consumption per agent action Covered in depth in the existing AI Workshack guide — same mechanism, credit-denominated interface

For scale on the API side, Anthropic reports that across enterprise Claude Code deployments the average cost is around $13 per developer per active day and $150 to $250 per developer per month, staying below $30 per active day for 90% of users. If your team is well outside that band, the habits below are where to look.

Habit 1: Clear between unrelated tasks (highest impact, lowest effort)

This is the single most valuable habit and almost nobody does it consistently. When you finish the auth bug and move to the CSS refactor, the auth conversation does not become free — it gets sent again on every message of the CSS work, forever. In Claude Code, /clear starts a fresh session and, from v2.1.211, resets the session cost counter to zero. If you might want the old session back, /rename it first so it is findable, then /resume later.

The important contrast is with compaction. Claude Code's docs are explicit that /compact reads the conversation it summarizes, so compacting a large context is itself a large request — whereas when you want a fresh start instead of continuity, /clear costs nothing. If the next task genuinely does not need the previous one, clearing is strictly cheaper than compacting.

Tradeoff: you lose everything the agent learned about your codebase in that session — the files it had already read, the conventions it had absorbed, the dead ends it knows to avoid. On a genuinely continuous piece of work, clearing means paying to rediscover all of it. Clear on task boundaries, not on time boundaries.

Habit 2: Scope the task before you start it

Vague requests trigger broad scanning. 'Improve this codebase' makes the agent read widely and speculatively, and every file it opens is carried for the rest of the session. 'Add input validation to the login function in auth.ts' lets it work with minimal file reads. Two habits compound this:

  • Use plan mode for anything non-trivial. The agent explores and proposes an approach for your approval before implementing, which prevents the expensive case: twenty turns of work in the wrong direction, all of which stays in context while it is undone.
  • Give verification targets up front — test cases, expected output, a screenshot of the bug. When the agent can check its own work it catches errors before you spend a round trip telling it what went wrong.

Course-correct early rather than late. Interrupting a wrong direction on turn three costs three turns; interrupting on turn twenty costs twenty, plus the inflated per-turn price of everything after. Claude Code exposes /rewind to restore both conversation and code to an earlier checkpoint — the cheap way to unwind a bad path. The alternative is arguing with the agent, and every round of that argument is permanently in context.

Tradeoff: tight scoping trades your time for the agent's tokens. On a subscription plan near a weekly limit that is obviously a good trade; on an unconstrained API budget it may not be.

Habit 3: Never let a large file or a verbose command land in context whole

This is where the biggest single-shot losses happen. A 40,000-token file read on turn three is 40,000 tokens on turn thirty as well. An unfiltered test run that prints every passing assertion, a full log tail, or a wide grep across a monorepo all land in context at full size and stay there. Three structural fixes, in ascending order of setup effort:

  1. Ask for the specific thing — 'show me the handler for POST /session in that file' rather than 'read that file', and failing tests rather than test output.
  2. Delegate verbose work to a subagent, so the output stays in its context and only a summary returns to yours. Claude Code's context-window walkthrough shows a research subagent doing several large file reads and returning roughly 420 tokens of summary.
  3. Preprocess with a hook. A PreToolUse hook can rewrite a command before it runs — grepping test output for failures, say — so the agent only sees matching lines. Anthropic's worked example reduces a log read from tens of thousands of tokens to hundreds.

For typed languages, code intelligence plugins are worth the setup: a 'go to definition' call replaces a grep plus several speculative file reads, and installed language servers report type errors automatically after edits.

Tradeoff: filtering means the agent sometimes lacks a detail it would have found by reading everything, then asks for it — costing an extra turn. Filters you write are also a maintenance surface. And the subagent route has its own cost, covered in the multi-agent article: the summary is cheap, but the subagent's own run is a full, separately-billed conversation.

Habit 4: Keep the always-on context small

Some tokens are in every request of every session regardless of what you are doing. Claude Code's published walkthrough of a representative session gives the shape: system prompt around 4,200 tokens, auto memory around 680, environment info around 280, MCP tool names around 120, skill descriptions around 450, user-level CLAUDE.md around 320, project CLAUDE.md around 1,800. You control the last four.

  • Keep your project instruction file lean — Anthropic's guidance is under 200 lines, essentials only. Detailed workflow instructions for PR reviews or database migrations sit in context even during unrelated work.
  • Move specialised instructions into skills, which load on demand when invoked rather than at session start.
  • Audit your MCP servers. Claude Code defers MCP tool definitions by default so only names enter context until a tool is used — but servers you never use are still listed. Disable them.
  • Prefer CLI tools over MCP servers where both exist: gh, aws, gcloud and sentry-cli add no per-tool listing at all.

Use the tool's own inspection command rather than guessing. Claude Code has /context for a live breakdown by category with optimisation suggestions, and /usage for token statistics; on paid plans /usage also attributes recent usage to skills, subagents, plugins and individual MCP servers, and flags behaviours such as long context or cache misses when one accounts for 10% or more of recent usage. That attribution view is the fastest way to discover a server you forgot about is eating a fifth of your allowance.

Tradeoff: a thinner project instruction file means the agent knows less about your conventions by default and you re-explain them more often. Moving instructions to skills means they only apply when something invokes the skill — a behaviour change, not just a cost change.

Habit 5: Fresh sessions versus long ones — the honest answer

'Always start fresh' is wrong, and so is 'never clear'. The right rule depends on whether the next work reuses the current context.

A long session is efficient when the agent has built up genuinely reusable knowledge of code you are still working on, and when your requests are close enough together that the prompt cache stays warm. Claude Code's cache lifetime is one hour on a subscription and five minutes by default on an API key or cloud provider, dropping to five minutes on a subscription once you are drawing on usage credits. A first message after a break longer than that misses the cache and reprocesses your full context at full price — which is why a long session resumed after lunch feels disproportionately expensive.

A long session is wasteful when it accumulates history that is no longer relevant — and it keeps costing even when you are not typing. Claude Code documents several ways an idle session keeps spending: scheduled tasks fire on their interval and send your full context each time, cross-session messages arrive as a new turn carrying full context, goal check-ins start a turn while background work is pending, and each active agent teammate consumes tokens until it exits. A practical rule: clear at task boundaries, compact within a task that has grown long, and close sessions you are not actively using rather than leaving them open overnight.

Compact deliberately, not reactively

Steer compaction rather than letting the automatic pass guess. Claude Code accepts instructions with the command — /compact focus on the auth bug fix — and lets you set how full the window gets before the automatic pass runs with a token count, for example /autocompact 500k. Standing compaction instructions can also live in CLAUDE.md.

Know what survives. In Claude Code the system prompt is untouched, project-root CLAUDE.md and auto memory are re-injected from disk, and up to five of the files the agent read or edited are re-read (most recently modified first) — but a file over 5,000 tokens comes back as a path reference rather than its content, and invoked skill bodies are capped at 5,000 tokens per skill and 25,000 total, oldest dropped first.

Tradeoff: compaction is lossy and it is not free — the summarisation request reads the entire conversation it is summarising. If continuity does not matter, clear instead.

Habit 6: Match the model to the job

Across the current Claude lineup as of August 2026, output tokens cost five times input, and the tiers run roughly 5:2:1 on input — claude-opus-5 at $5 in / $25 out per million, claude-sonnet-5 at $2 / $10, claude-haiku-4-5 at $1 / $5. Anthropic's own guidance is that Sonnet handles most coding tasks well and costs less than Opus, and to reserve the Opus tier for complex architectural decisions or multi-step reasoning. Two things to know before you downgrade everything:

  • Anthropic's cost documentation identifies unexpectedly high spend on API and cloud-provider plans as usually tracing back to two causes: long sessions that were never cleared, or the Opus tier left as the default model. Both are one-setting fixes.
  • On a subscription plan, a session or weekly limit is shared across all models, so switching models with /model will not restore access. A model-specific limit message ('you've hit your Opus limit') is different — switching to a model outside that family does keep you working.

For narrow mechanical work — running a test suite, reformatting, fetching and summarising docs, scanning for a pattern — the cheapest tier is usually sufficient and can be pinned per subagent rather than per session. A Claude Code subagent definition takes a model field in frontmatter accepting a family alias (sonnet, opus, haiku, fable), a full model ID, or inherit, which is the default.

---
name: test-runner
description: Runs the test suite and reports only failures
model: haiku
---

Tradeoff: this is the technique most likely to backfire. A cheaper model that takes twelve turns where a stronger one takes four is a net loss, because each turn is billed the price of the whole conversation so far. Downgrade for narrow, well-specified, mechanical work — not for tasks that need judgement about your architecture.

Habit 7: Turn reasoning down for work that does not need it

Extended thinking is on by default in Claude Code because it materially improves performance on complex planning and reasoning. It is also billed as output tokens — the expensive side — and the default budget can reach tens of thousands of tokens per request. On an agentic loop that reasoning happens every turn.

The current control is the effort level, adjustable with /effort or in /model. On models that still use a fixed thinking budget rather than adaptive reasoning, a MAX_THINKING_TOKENS environment variable sets an explicit ceiling; adaptive-reasoning models ignore nonzero budgets, so effort is the lever there. Thinking can also be disabled entirely in /config on models that permit it.

Tradeoff: this is a direct quality lever. Lower effort produces fewer, more consolidated tool calls and terser output — exactly right for mechanical work and exactly wrong for debugging something subtle. Cheap reasoning that produces ten extra turns of flailing costs more than the thinking you saved.

Habit 8: Watch for cache misses, not just token counts

A cache read costs 0.1x the base input rate, so a well-cached session pays roughly a tenth for its re-sent history. Losing the cache does not change your token count — it changes what those tokens cost by a factor of ten, which makes cache misses invisible in any dashboard that only counts tokens.

The common causes are breaks longer than the cache lifetime, and anything that changes the cached prefix. The prefix hierarchy is tools, then system, then messages, so changing a tool definition invalidates all three: enabling or disabling an MCP server mid-session, or editing CLAUDE.md while working, both throw the cache away. Claude Code's /usage flags cache misses when they account for 10% or more of recent usage; on an API integration, check the cache read token count in the response usage directly.

The short version

  1. Clear between unrelated tasks — free, and the highest-impact habit available.
  2. Scope tightly and plan before implementing; wrong directions are the expensive failure mode.
  3. Keep big files and verbose output out of context — ask narrowly, delegate, or filter with a hook.
  4. Trim the always-on context: lean project instructions, on-demand skills, no unused MCP servers.
  5. Compact deliberately with instructions; clear instead when continuity does not matter.
  6. Match the model to the job, and check your default is not the top tier by accident.
  7. Lower reasoning effort for mechanical work only.
  8. Treat cache misses as a first-class cost signal.
If you only change one thing this week: run your tool's context-inspection command mid-session on a normal working day. Most people find one item — a forgotten MCP server, a bloated instruction file, a habitual full-file read — that accounts for a surprising share of everything they spend.
  • Why AI agents burn tokens — the underlying mechanic these habits are working against.
  • Multi-agent cost control — when delegating to subagents saves money and when it multiplies the bill.
  • Replit Agent credit consumption: why costs are unpredictable and what to do about it.
  • LiteLLM: one API to route all your LLMs — routing coding-agent traffic through a proxy for per-key spend tracking.
  • Langfuse 101 — making token distribution observable rather than inferred.
Verification note: command names and behaviours (/clear, /compact, /autocompact, /context, /usage, /effort, /model, /rewind, /rename, /resume, MAX_THINKING_TOKENS, subagent model frontmatter) were checked against Claude Code documentation in August 2026. Cursor and Copilot billing mechanics were checked against their own current pricing documentation. Where a practice applies across tools but the syntax does not, it is described without inventing a command.