Why output costs 5x input, what max_tokens actually does (it is not a budget), which concision instructions work, and how to measure the difference.

Most cost-reduction effort goes into the prompt. That is the cheap half. On every current Claude model, output tokens are priced at exactly five times input tokens — Opus 4.8 at $5 in and $25 out, Sonnet 5 at $2 and $10, Haiku 4.5 at $1 and $5. A hundred words the model writes cost as much as five hundred words you send.

That ratio changes where the leverage is. A verbose response is not a style problem, it is a line item. And on reasoning models the problem is worse than it looks, because a substantial and completely invisible portion of your output tokens is spent before the model writes a single word you will ever see.

Prices and parameter behaviour in this article were verified against Anthropic and OpenAI documentation as of August 2026. Verify current pricing and API signatures before relying on them. The 5:1 output-to-input ratio is more durable than the absolute figures, but it too can change.

Reasoning tokens are output tokens

This is the single most under-appreciated line in anyone's bill. On OpenAI's reasoning models, tokens spent on internal reasoning are billed as output tokens, are not returned to you, and count toward your output limit. They appear in the usage object under `output_tokens_details.reasoning_tokens`.

{
  "usage": {
    "input_tokens": 75,
    "input_tokens_details": { "cached_tokens": 0 },
    "output_tokens": 1186,
    "output_tokens_details": { "reasoning_tokens": 1024 },
    "total_tokens": 1261
  }
}

In that example, 1,024 of 1,186 output tokens — 86% — are reasoning you never see. If you were estimating cost from the length of the visible answer, you were off by a factor of seven. The same principle applies on Anthropic models, where thinking tokens are billed as output.

The lever for this is the effort control, and it is enforced by the platform rather than requested in prose. On Anthropic, `output_config.effort` accepts `low`, `medium`, `high`, `xhigh` and `max`, defaulting to `high` on Opus 4.8 and Sonnet 5. On OpenAI, `reasoning.effort` accepts values from `none` and `minimal` through to `max`. Sweeping effort downward on a task that does not need deep reasoning is the highest-return single-parameter change available on output cost.

Before optimising anything else about your output, log `reasoning_tokens` (OpenAI) or compare `output_tokens` against the visible text length (Anthropic) for a day. If invisible reasoning dominates, effort tuning is your whole answer and no amount of concision instruction will touch it.

max_tokens is a circuit breaker, not a budget

Almost everyone misreads this parameter. `max_tokens` on Anthropic (required, and `max_output_tokens` on OpenAI) does not tell the model to be brief. The model is not aware of the limit and does not pace itself against it. It generates as it would have generated, and the API cuts it off when the ceiling is hit.

Three consequences, all of them expensive:

  1. **You pay for everything generated up to the cut.** Truncation does not refund you. A 4,000-token response capped at 500 tokens still bills what it produced up to the cap.
  2. **Truncated output is usually unusable, so it gets retried.** A retry is a second full request. Setting `max_tokens` too low reliably makes a workload more expensive, not less.
  3. **On reasoning models the cap can consume itself.** OpenAI's documentation is explicit that `max_output_tokens` limits all generated tokens including non-visible ones, and warns that reasoning "might occur before any visible output tokens are produced, meaning you could incur costs for input and reasoning tokens without receiving a visible response." You can pay in full and get an empty answer. Their guidance is to reserve at least 25,000 tokens for reasoning and output when starting out with these models.

So set `max_tokens` generously, as a blast-radius limit against a runaway generation, and detect when you hit it rather than treating it as a cost control.

import anthropic
 
client = anthropic.Anthropic()
 
response = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=4096,  # headroom, not a budget
    messages=[{"role": "user", "content": task}],
)
 
if response.stop_reason == "max_tokens":
    # This is a cost incident: you paid for 4096 tokens and got a fragment.
    # Alert on it. Do not silently retry with a bigger cap in a loop.
    log_truncation(model="claude-sonnet-5", output_tokens=response.usage.output_tokens)

On OpenAI the equivalent signal is a response with `status` set to `incomplete` and `incomplete_details.reason` equal to `max_output_tokens`. Either way, treat a rising truncation rate as a regression — it is a direct measure of money spent on unusable output.

Structured outputs: the strongest control you have

If you want an output length guarantee that is actually enforced, use constrained decoding. A JSON schema does not ask the model to be brief; it makes anything outside the schema impossible to generate. No "Sure! Here's the JSON:" preamble, no markdown fence, no closing paragraph explaining what it just did.

import anthropic
from pydantic import BaseModel
from typing import Literal
 
client = anthropic.Anthropic()
 
class TicketTriage(BaseModel):
    category: Literal["billing", "technical", "account", "other"]
    urgency: Literal["low", "medium", "high"]
    summary: str
 
response = client.messages.parse(
    model="claude-haiku-4-5",
    max_tokens=512,
    messages=[{"role": "user", "content": ticket_text}],
    output_format=TicketTriage,
)
 
triage = response.parsed_output      # typed, guaranteed schema-valid
print(response.usage.output_tokens)  # compare this against the free-text baseline

The raw form uses `output_config` on `messages.create()` if you would rather build the schema by hand:

response = client.messages.create(
    model="claude-haiku-4-5",
    max_tokens=512,
    messages=[{"role": "user", "content": ticket_text}],
    output_config={
        "format": {
            "type": "json_schema",
            "schema": {
                "type": "object",
                "properties": {
                    "category": {"type": "string", "enum": ["billing", "technical", "account", "other"]},
                    "urgency": {"type": "string", "enum": ["low", "medium", "high"]},
                    "summary": {"type": "string"},
                },
                "required": ["category", "urgency", "summary"],
                "additionalProperties": False,
            },
        }
    },
)

OpenAI's equivalent lives under `text.format` with `strict` set to true, and the SDK offers `client.responses.parse(..., text_format=YourModel)` returning `response.output_parsed`. The mechanism is the same on both platforms: the decoder is constrained by a grammar, so schema compliance is structural rather than instructed.

Two design choices inside the schema

  • **Enums are free brevity.** A `Literal` or `enum` field emits a handful of tokens, guaranteed. Every classification field you can express as an enum instead of a free string is an output reduction you never have to re-verify.
  • **Field names are emitted on every single response.** `{"recommended_next_action": ...}` costs more per call than `{"next_action": ...}`. This is a genuine lever at high volume — and a genuine readability cost. Trim obviously redundant naming; do not compress your schema into initials to save a rounding error.
Structured outputs are not free. Per Anthropic's documentation they inject an additional system prompt describing the output format, which increases your input token count; the first request against a new schema pays a grammar-compilation latency cost (compiled grammars are cached for 24 hours); and changing `output_config.format` invalidates the prompt cache for that conversation. Note also that string-length constraints such as `maxLength` are not among the supported schema keywords — you cannot cap a field's length with the schema itself.

Stop sequences: narrow, but exact

`stop_sequences` takes an array of strings and halts generation the moment one appears. The response comes back with `stop_reason` set to `stop_sequence` and a `stop_sequence` field containing which one matched.

response = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    stop_sequences=["\n\nNote:", "\n\nWould you like"],
    messages=[{"role": "user", "content": task}],
)
 
if response.stop_reason == "stop_sequence":
    print("cut at:", response.stop_sequence)

This is a scalpel, not a general concision tool. It works when you have observed a specific recurring tail — a trailing "Let me know if you'd like me to expand on any of this" that shows up on 40% of responses — and you want it gone deterministically. It does nothing about verbosity spread evenly through a response, and if the model never emits your sequence you get no benefit at all. Note also that the matched text is not returned, so anything you stop on is content you have decided you never want.

Concision instructions: what works and what does not

Prompt-level concision is the weakest control here because it is a request rather than a constraint. But it is also the only one that works on free-text output, so it is worth knowing which forms carry weight.

Instruction Works? Why
"Be concise." / "Be brief." Barely No target, no definition. Interpretation drifts across prompts and models.
"Keep your response under 200 tokens." No Models cannot count their own tokens. The unit is unobservable to the writer.
"Answer in at most 3 bullets, each under 15 words." Yes Countable units the model can track while generating.
"Respond with exactly one line: `<category>: <reason>`" Yes A format contract, not an adjective.
"No preamble. Do not restate the question. No closing summary." Yes Names the specific padding you actually observed.
A worked example of the target length Yes — strongest prompt-level lever Demonstration beats description; the model matches the shape it was shown.
"Do not be verbose." Barely Negations of vague adjectives are the weakest instruction form available.

The pattern is clear enough to state as a rule: constrain in units the model can count while writing — sentences, bullets, lines, fields — and never in units it cannot, such as tokens, characters or bytes. And prefer showing over telling. One example of a correctly-sized response outperforms three sentences of instruction about length, consistently.

OpenAI also exposes `text.verbosity` with values `low`, `medium` and `high`, which is a platform-level control rather than a prompt instruction and therefore more reliable than any wording. If you are on those models, set it before you write a word about brevity.

Write your concision rules against padding you have actually measured, not padding you imagine. Pull fifty real responses, mark the parts nobody reads, and write instructions naming those specific patterns. Generic brevity language written from intuition tends to shorten the useful content and leave the boilerplate intact.

Measuring the effect properly

Output tokens cannot be predicted before a call. The token-counting endpoint — free, and rate-limited separately from message creation at 2,000 to 8,000 requests per minute depending on your usage tier — counts input only. Output measurement is necessarily after the fact, which means logging.

import statistics
import anthropic
 
client = anthropic.Anthropic()
 
def measure(prompt_variant: str, eval_inputs: list[str]) -> dict:
    outputs = []
    for text in eval_inputs:
        r = client.messages.create(
            model="claude-sonnet-5",
            max_tokens=2048,
            system=prompt_variant,
            messages=[{"role": "user", "content": text}],
        )
        outputs.append(r.usage.output_tokens)
 
    outputs.sort()
    return {
        "median": statistics.median(outputs),
        "mean": statistics.mean(outputs),
        # The tail is where the money is. A handful of runaway responses can
        # dominate total spend while leaving the mean almost unchanged.
        "p95": outputs[int(len(outputs) * 0.95) - 1],
        "max": outputs[-1],
    }
 
baseline = measure(CURRENT_SYSTEM_PROMPT, EVAL_INPUTS)
variant = measure(CONCISE_SYSTEM_PROMPT, EVAL_INPUTS)
 
print(f"median {baseline['median']} -> {variant['median']}")
print(f"p95    {baseline['p95']} -> {variant['p95']}")

Report the distribution, not the average. A prompt change that cuts the median by 10% but the p95 by 60% is a much bigger win than the mean suggests, because total spend is driven by the long responses. Conversely a change that improves the mean while leaving the tail intact has not fixed your actual cost problem.

  • **Always pair the token delta with a quality score.** Shorter is trivially achievable and frequently worse. A saving reported without a quality number is not a result.
  • **Fix the eval set and reuse it.** Comparing this week's production traffic against last week's measures input drift, not your change.
  • **Tag every request with the prompt version.** Without it you cannot attribute a cost movement to a specific change, and prompt changes are the most frequent uncontrolled variable in an LLM system.
  • **Never compare token counts across model generations.** Anthropic's documentation notes that Claude 4.7 and later models use a newer tokenizer producing roughly 30% more tokens for the same text, and advises recounting against the model you actually plan to use.

For continuous tracking rather than one-off experiments, push output tokens, prompt version and quality score into a tracing layer. The AI Workshack guides to LangFuse and the LangSmith comparison cover wiring that up so output tokens by prompt version becomes a chart you can watch, and the LiteLLM guide covers gateway-level cost attribution when you want the same numbers in currency across multiple providers.

Every control, and what it costs you

Control Enforcement What you give up
Effort / reasoning effort Platform-enforced Reasoning depth on hard tasks. The highest-leverage single change.
Structured outputs Decoder-enforced (grammar) Extra input tokens for the injected format prompt, first-call compile latency, prompt cache invalidation on schema change, schema-keyword limits.
`text.verbosity` (OpenAI) Platform-enforced Detail. Reliable, low effort.
Stop sequences Hard cut at the string Only works on a predictable tail; the matched text is discarded.
Concision instructions Model discretion Unreliable, drifts across prompt edits and model upgrades. Can shorten substance rather than padding.
`max_tokens` / `max_output_tokens` Hard cut, model unaware Not a cost control. Truncated output is paid for and usually retried, doubling the cost.

The order to do this in

  1. Log `output_tokens` — and `reasoning_tokens` if you are on OpenAI reasoning models — for one week, tagged by endpoint and prompt version. Find the p95, not the mean.
  2. If invisible reasoning dominates, sweep the effort parameter down one level at a time against your eval set. Stop when quality moves. This is usually the whole win.
  3. Convert anything with a fixed shape to structured outputs. Classification, extraction, routing and scoring all qualify, and the schema enforces brevity permanently rather than per-prompt.
  4. Look at fifty real free-text responses and write concision rules that name the padding you actually see, with a worked example of the target length.
  5. Add stop sequences only for a specific recurring tail you have measured.
  6. Set `max_tokens` generously and alert on `stop_reason == "max_tokens"` as a cost incident.
  7. Re-measure with the same frozen eval set, report median and p95 alongside a quality score, and only then call it a saving.
Pricing, parameter names and behaviour on this page were verified against Anthropic and OpenAI documentation in August 2026. API surfaces and rates both change — re-check the current docs before you build on any specific figure or signature here.