Tagging strategy, computing cost from usage fields, unit economics, and chargeback that survives an audit — including the awkward cases of cached and batched calls.

The console says $48,000 this month. Finance wants to know whether the enterprise pilot is profitable. Product wants to know whether the new document-summary feature is worth keeping. Support wants to know which customer is generating the load. The console cannot answer any of those questions, and no amount of staring at it will change that.

Attribution is not a reporting problem. It is a data-modelling problem that has to be solved at the moment the request is made, because the information you need — which user, which feature, which tenant — exists in your application and nowhere else. By the time the call reaches the provider, it is gone.

All prices in this article are as of August 2026 and are quoted only to make the arithmetic concrete. Verify current pricing against the provider's own pricing page before you put a number in a spreadsheet, a budget, or a customer contract.

Step 1: decide the dimensions before you write any code

Attribution quality is capped by the dimension set you chose on day one. Adding a dimension later means every historical row lacks it, and backfilling from logs is a project. Pick deliberately.

Dimension Why you need it Cardinality Watch out for
tenant_id The unit finance bills and the unit that churns Low to medium Internal and trial tenants must be flagged, not mixed in
user_id Power-user detection, abuse, per-seat economics High Must be pseudonymous; never an email address
feature The only way to kill an unprofitable feature Low Needs a controlled vocabulary or it becomes free text
environment Keeps dev, CI and load tests out of customer numbers Very low CI is often the single largest 'customer' by call count
release / prompt_version Attributes a cost regression to a deploy Medium Must change when the prompt changes, not only when code ships
session_id / conversation_id Agent runs and multi-turn cost High Needed to catch runaway loops
attempt Separates retried spend from first-try spend Very low Without it, retries silently inflate per-user cost
Controlled vocabulary for feature is worth enforcing in code — an enum, not a string. Free-text feature tags degrade into forty spellings of the same thing within a quarter, and the resulting dashboard is unusable in exactly the month you need it.

Step 2: attach at the edge, propagate down

The single most common attribution failure is tagging at the call site. Somebody adds metadata to the three obvious LLM calls, then a fourth is added in a background job six months later with no tags, and now 12% of spend is unattributed and nobody notices because unattributed spend does not appear on a per-tenant dashboard.

The fix is structural: build the attribution context once, at the request boundary, and have every LLM call read it from context rather than receive it as an argument. A call that forgets its tags should be impossible, not merely discouraged.

from contextvars import ContextVar
from dataclasses import dataclass, asdict, field
 
@dataclass(frozen=True)
class CostContext:
    tenant_id: str
    user_hash: str          # pseudonymous, never raw PII
    feature: str            # from a controlled enum
    environment: str        # prod | staging | ci
    release: str
    request_id: str
    session_id: str | None = None
    attempt: int = 1
 
    def as_tags(self) -> dict[str, str]:
        return {k: str(v) for k, v in asdict(self).items() if v is not None}
 
_ctx: ContextVar[CostContext | None] = ContextVar("cost_ctx", default=None)
 
def set_cost_context(ctx: CostContext) -> None:
    _ctx.set(ctx)
 
def current_cost_context() -> CostContext:
    ctx = _ctx.get()
    if ctx is None:
        # Fail loudly in tests and CI; fall back to 'unattributed' in prod.
        raise RuntimeError("LLM call made outside a cost context")
    return ctx
 

Set it once in middleware, and for background jobs set it in the job wrapper from whatever enqueued the work. Then each transport renders the same context into whatever shape that vendor accepts.

import anthropic, hashlib
 
client = anthropic.Anthropic()
 
def call_model(prompt: str, model: str = "claude-sonnet-5"):
    ctx = current_cost_context()
    return client.messages.create(
        model=model,
        max_tokens=1024,
        messages=[{"role": "user", "content": prompt}],
        # The Messages API metadata object accepts user_id only, max 512 chars.
        # It must be an opaque identifier - no names, emails or phone numbers.
        metadata={"user_id": ctx.user_hash},
    )
 
Anthropic's metadata.user_id is deliberately narrow: one optional opaque string, up to 512 characters, used to help detect abuse. It is not an attribution system and it will not appear broken out in your cost report. Your tenant, feature and release dimensions have to live in your own telemetry, not in the provider request.

For the other transports, the shapes are:

  • Helicone: HTTP headers. Helicone-User-Id for the reserved user dimension, Helicone-Property-Feature, Helicone-Property-Tenant and so on for the rest, Helicone-Session-Id to group an agent run.
  • Portkey: an x-portkey-metadata header carrying a JSON object, or the metadata parameter in its SDKs. Values must be strings of at most 128 characters, and the _user key is the one that powers user-level analytics.
  • LangFuse: propagate_attributes() in the current Python SDK sets user_id, session_id, tags and metadata for everything downstream in the context — see our existing LangFuse articles for the full instrumentation walkthrough.
  • OpenTelemetry: span attributes under your own namespace, alongside the standard gen_ai.* usage attributes.

Step 3: compute cost per request from the usage fields

Every response carries the raw material. For Anthropic's Messages API the usage object contains input_tokens (only the tokens after the last cache breakpoint), cache_creation_input_tokens, cache_read_input_tokens and output_tokens, and when the one-hour cache TTL is in use a cache_creation object breaking writes into ephemeral_5m_input_tokens and ephemeral_1h_input_tokens. Total input tokens processed is the sum of the three input figures — a detail that catches people out, because input_tokens on its own understates the request badly when caching is on.

Prices as of August 2026, quoted only to make the code below concrete — verify current pricing before use:

Model Input / MTok 5m cache write 1h cache write Cache read Output / MTok
claude-opus-5 $5.00 $6.25 $10.00 $0.50 $25.00
claude-sonnet-5 $2.00 $2.50 $4.00 $0.20 $10.00
claude-haiku-4-5 $1.00 $1.25 $2.00 $0.10 $5.00

The structure is stable even when the numbers are not: a five-minute cache write is 1.25x the base input rate, a one-hour write is 2x, and a cache read is 0.1x. Batch processing applies a 50% discount to both input and output, and it stacks with caching.

from decimal import Decimal
 
PRICE_BOOK_VERSION = "2026-08-04"
 
# USD per million tokens. Verify against the provider pricing page.
PRICES = {
    "claude-opus-5":            {"in": "5.00", "w5m": "6.25", "w1h": "10.00", "read": "0.50", "out": "25.00"},
    "claude-sonnet-5":            {"in": "2.00", "w5m": "2.50", "w1h": "4.00",  "read": "0.20", "out": "10.00"},
    "claude-haiku-4-5":  {"in": "1.00", "w5m": "1.25", "w1h": "2.00",  "read": "0.10", "out": "5.00"},
}
MILLION = Decimal(1_000_000)
 
def price_request(model: str, usage, batch: bool = False) -> dict:
    rates = PRICES[model]
    creation = getattr(usage, "cache_creation", None)
    w5m = getattr(creation, "ephemeral_5m_input_tokens", None) if creation else None
    w1h = getattr(creation, "ephemeral_1h_input_tokens", None) if creation else None
    if w5m is None and w1h is None:
        # No TTL breakdown available: assume the 5-minute rate.
        w5m = getattr(usage, "cache_creation_input_tokens", 0) or 0
        w1h = 0
 
    buckets = {
        "input_uncached": (usage.input_tokens, rates["in"]),
        "cache_write_5m": (w5m or 0, rates["w5m"]),
        "cache_write_1h": (w1h or 0, rates["w1h"]),
        "cache_read":     (getattr(usage, "cache_read_input_tokens", 0) or 0, rates["read"]),
        "output":         (usage.output_tokens, rates["out"]),
    }
    multiplier = Decimal("0.5") if batch else Decimal("1")
    cost_details = {
        k: (Decimal(tokens) * Decimal(rate) / MILLION * multiplier)
        for k, (tokens, rate) in buckets.items()
    }
    return {
        "usage_details": {k: v[0] for k, v in buckets.items()},
        "cost_details": {k: float(v) for k, v in cost_details.items()},
        "cost_usd": float(sum(cost_details.values())),
        "price_book_version": PRICE_BOOK_VERSION,
    }
 

Note the bucket discipline: each token is counted in exactly one bucket, which is the same rule LangFuse enforces on its usage_details and cost_details fields. Overlapping buckets — counting cached tokens in both input and cache_read — is the single most common way a homegrown cost calculator ends up 30% above the invoice.

Where your computed number will drift from the invoice

It will drift. Name the causes up front so the monthly variance is a known quantity rather than a fire drill:

  • Negotiated or promotional rates that your hard-coded price book does not know about.
  • Server-side tool charges billed separately from tokens — web search, for example, is priced per search rather than per token.
  • Container or runtime charges, such as code execution billed by the hour rather than by the token.
  • Regional or data-residency multipliers applied on top of base rates.
  • Premium service modes priced above the standard rate for the same model.
  • Rounding. Per-request decimals accumulate differently to provider-side aggregation.
Track computed-versus-invoiced as a first-class monthly metric with a tolerance band, say 2%. A drift that stays inside the band is fine. A drift that jumps is telling you a pricing change or a new billable feature landed without you noticing.

Step 4: build the unit economics

Attribution is the input; unit economics is the output. Three denominators cover most of what a business actually asks.

Cost per active user

Total attributable spend for a period divided by users who made at least one AI-backed action. The trap is the denominator: monthly active users of your product is the wrong number, because most of them never touched the AI feature. Use AI-active users, and report the ratio of AI-active to total active alongside it — that ratio is what tells you whether cost scales with adoption or with headcount.

Cost per feature, per invocation

Feature spend divided by feature invocations. This is the number that kills features. A summariser at $0.004 per invocation and a research agent at $0.90 per invocation are different businesses, and the aggregate dashboard shows them as one line.

Cost per resolved task

The most useful and the most expensive to build, because it requires an outcome signal: the ticket was closed, the draft was accepted, the answer was not followed by a rephrase. Without it, you are optimising cost per attempt, which rewards a cheap model that fails twice as often.

The tradeoff is honest and unavoidable: defining resolution is product work, not platform work, and it needs an evaluation pipeline behind it. Our RAG evaluation articles cover the measurement side. What matters here is that cost per attempt and cost per resolution can move in opposite directions, and only one of them is the business metric.

Scenario Cost per attempt Success rate Cost per resolution
Larger model, single pass $0.090 88% $0.102
Smaller model, single pass $0.018 61% $0.030
Smaller model, verify and retry once on failure $0.029 84% $0.035

Illustrative figures only — the point is the shape. The cheap-per-attempt option is genuinely cheaper per resolution here, but the gap narrows sharply once retries are counted, and it inverts entirely in domains where a failure has a downstream human cost. Measure it for your workload; do not assume it.

Step 5: showback and chargeback that survive an audit

Showback Chargeback
What it is Reporting spend to the team that caused it Moving the money to their budget
Accuracy bar Directionally right Defensible line by line
Disputes Rare Guaranteed
Prerequisite Attribution Attribution plus a written allocation policy
Good first step Yes Only after two quiet quarters of showback

Start with showback. Teams change behaviour when they can see their number, and you will discover every flaw in your attribution during the showback period rather than during a budget dispute.

An attribution ledger that survives audit needs five properties:

  1. Reproducible. Same inputs, same output. That means an immutable price book keyed by version, not a table somebody edits in place.
  2. Reconcilable. Every period's total ties to the provider's cost report, with the residual explained rather than silently absorbed.
  3. Append-only. Corrections are new adjusting rows with a reason, not updates to historical rows.
  4. Complete. There is an explicit unattributed bucket, and its size is a tracked metric. A ledger with no unattributed bucket is a ledger that is quietly lying.
  5. Retained. The raw usage rows, not just the aggregates, for as long as your finance policy requires.

Allocating the residual

Some spend genuinely belongs to no tenant: evaluation runs, warm-up calls, internal tooling, failed requests, platform overhead. You have three options, each with a cost:

  • Leave it in a platform bucket. Cleanest and most honest. Tradeoff: someone owns a budget line that only grows, and it becomes politically hard to defend.
  • Allocate proportionally to attributed usage. Simple and defensible. Tradeoff: it penalises large well-behaved tenants for overhead they did not cause.
  • Allocate as a flat per-tenant platform fee. Matches how most SaaS pricing already works. Tradeoff: it materially over-charges small tenants and will be disputed by them first.

Whichever you pick, write it down in one page before the first chargeback cycle. The policy being boring and public is what makes it survivable; a policy invented during a dispute never does.

The awkward cases: cache and batch

Cached calls

Caching breaks the assumption that the beneficiary and the payer are the same request. Whoever triggers a cache write pays roughly 1.25x the base input rate for tokens that everyone afterwards reads at 0.1x. Charge naively and the first tenant to hit a shared prompt each cache window subsidises the rest.

Three workable policies:

Policy How it works Tradeoff
Charge the writer Cache write cost lands on the triggering request Simple and ties exactly to the invoice, but visibly unfair on shared prompts
Socialise writes Cache writes go to a platform pool, reads charged as incurred Fair, but tenant totals no longer sum to the invoice without the pool line
Charge list price, bank the saving Every tenant is charged as if uncached; the delta is a central credit Excellent for showing avoided cost, but tenant charges exceed actual spend - never use this for real chargeback

If the cached content is tenant-specific rather than shared — a tenant's own documents, their own system prompt — charge the writer. There is no fairness problem, because the tenant is the only beneficiary.

Batched calls

Batch attribution fails for a structural reason: results come back detached from the request that created them, in a different order, possibly hours later. There is no context variable alive to read.

The fix is to carry attribution in the batch record itself. Anthropic's Message Batches API gives every request a custom_id, results are explicitly not guaranteed to be in submission order, and custom_id is the documented way to match a result to its request. Encode a durable reference there — not the whole context, which will not fit and should not be in a provider-side field, but a key into your own job table.

# At enqueue: persist the full attribution context keyed by a job id,
# and put only that key in custom_id.
job_id = persist_attribution(current_cost_context(), item_id=item.id)
requests.append({
    "custom_id": job_id,          # e.g. "job_01H9X..." - your key, not PII
    "params": {
        "model": "claude-haiku-4-5",
        "max_tokens": 512,
        "messages": [{"role": "user", "content": item.text}],
    },
})
 
# At result time: look the context back up, and price with batch=True
# so the 50% discount is reflected in the tenant's ledger row.
for result in stream_batch_results(batch_id):
    ctx = load_attribution(result.custom_id)
    priced = price_request(model, result.result.message.usage, batch=True)
    write_ledger_row(ctx, priced, service_tier="batch")
 
Do not put tenant names, user identifiers or anything else recognisable in custom_id. It is a provider-side field carried through logs and result files. Use an opaque key that only means something inside your database.

Retries, fallbacks and failed calls

Three more leaks worth closing. A retry is billable and must be attributed to the original request with an incremented attempt, or per-user cost inflates without any dimension explaining why. A fallback to a different model must record the model actually used, not the model requested. And a call that errors after the provider processed the input is still billable input — attribute it, tag it as failed, and watch that bucket, because a rising failed-spend line is one of the cleanest early signals of a broken deploy.

A verification checklist

  1. Pick a random hour. Does the sum of your ledger rows for that hour match the native usage API for the same window, within tolerance?
  2. What percentage of spend is in the unattributed bucket? If you do not know, that is the answer.
  3. Take your largest tenant. Can you produce a line-item explanation of their number in under ten minutes?
  4. Turn caching off in staging. Does your per-request cost change in the direction and magnitude you predicted?
  5. Submit a batch. Does every result land on a tenant, at the discounted rate?
  6. Force a retry storm in staging. Does the cost land on the right request with attempt greater than one?

Once attribution is trustworthy, it becomes the substrate for control: budgets, alerts and anomaly detection, which is the subject of the next article in this track.