Native usage APIs, observability platforms, gateway metering and OpenTelemetry — what each layer actually gives you, what it costs you, and how to pick without locking yourself in.

Every team that ships on an LLM eventually asks the same three questions in the same order: what did we spend, who spent it, and is that number about to get worse. The first question has an easy answer. The second and third do not, and the reason is almost always that the team instrumented the wrong layer.

There are four places you can meter LLM spend, and they are not substitutes for each other. They sit at different points in the request path, they see different things, and they fail in different ways. This article walks each one, says plainly what it does not give you, and ends with a decision table you can actually use in a design review.

All prices in this article are as of August 2026 and are quoted only to make the arithmetic concrete. Verify current pricing against the provider's own pricing page before you put a number in a spreadsheet, a budget, or a customer contract.

The four layers, in one table

Layer Where it sits Best at Blind to
Native provider usage/cost APIs Provider's billing system Ground truth for reconciliation Per-user, per-feature, per-request detail
Observability platform (SDK or proxy) Your app, or in front of it Traces, prompts, evals, attribution Spend on calls it never saw
Gateway metering A proxy every call goes through Enforcement: budgets, keys, routing Anything that bypasses the gateway
OpenTelemetry GenAI conventions Your instrumentation code Vendor-neutral, joins to the rest of your stack Cost — the spec has no money attribute

The short version: the native APIs tell you the truth, the observability platform tells you the story, the gateway lets you do something about it, and OpenTelemetry keeps you from being owned by any of them. Most mature setups run at least two.

Layer 1: native provider usage and cost APIs

Both major providers now expose admin-level usage and cost endpoints. They are underused, and they are the only numbers that will ever match your invoice.

Anthropic

Two endpoints, both requiring an Admin API key (prefixed sk-ant-admin01-, distinct from a normal API key, and unavailable on individual accounts):

  • GET /v1/organizations/usage_report/messages — token consumption. Time buckets of 1m, 1h or 1d. You can filter and group by model, workspace, API key, service tier, context window and inference geo. Bucket count limits differ by granularity: 1m defaults to 60 buckets and caps at 1,440; 1h defaults to 24 and caps at 168; 1d defaults to 7 and caps at 31.
  • GET /v1/organizations/cost_report — cost in USD, daily granularity only, groupable by workspace and description.

Data typically appears within about five minutes of the request completing, and the documented sustainable polling frequency is once per minute. Both endpoints paginate with has_more and next_page.

curl "https://api.anthropic.com/v1/organizations/usage_report/messages?\
starting_at=2026-08-01T00:00:00Z&\
ending_at=2026-08-08T00:00:00Z&\
group_by[]=model&\
group_by[]=api_key_id&\
bucket_width=1d" \
  -H "anthropic-version: 2023-06-01" \
  -H "x-api-key: $ANTHROPIC_ADMIN_KEY"

OpenAI

The same shape, different nouns. GET /v1/organization/usage/completions accepts bucket_width of 1m, 1h or 1d and groups by project_id, user_id, api_key_id, model, batch and service_tier. GET /v1/organization/costs is daily only and groups by project_id, line_item and api_key_id. Both take start_time and end_time as Unix seconds and require an admin key.

A trap worth naming: the user_id dimension in OpenAI's usage API refers to a member of your OpenAI organization — the human who owns the key — not the end user of your product. It will never tell you which of your customers is expensive.

What the native APIs genuinely do not give you

  • Per-end-user attribution. The finest grain available is the API key, project or workspace.
  • Per-feature attribution. Your summariser and your chat assistant look identical if they share a key.
  • Per-request detail. You get buckets, not rows. You cannot open the one call that cost $4.
  • Prompts, responses or latency. No correlation between spend and quality, errors or user experience.
  • A cross-provider view. Two providers means two schemas, two admin keys and a join you write yourself.

The common workaround is a key per tenant or per feature. It works, and it is the cheapest possible attribution mechanism, but be honest about the tradeoff: key sprawl, a rotation burden that grows linearly with tenants, and provider-side performance limits when you break costs down across a large number of keys. Anthropic's own documentation steers high-cardinality per-user reporting away from the usage API for exactly this reason. Treat key-per-tenant as viable up to tens of keys, not thousands.

Whatever else you build, wire the native cost API into a monthly reconciliation job. Any cost you compute yourself is an estimate; this is the number your finance team will be invoiced for. The gap between them is a metric in its own right.

Layer 2: LLM observability platforms

This is the layer that answers "who and why". AI Workshack already has dedicated guides to LangFuse and LangSmith, including a head-to-head comparison, so this section covers the two the existing articles do not: Helicone and Portkey. Read those pieces for the LangFuse and LangSmith detail — the selection criteria below apply to all four.

Helicone

Helicone's default integration is a proxy: you change your base URL, and attribution rides in HTTP headers. Helicone-User-Id is a reserved header that drives per-user cost analytics; Helicone-Property-<Name> attaches arbitrary custom properties; Helicone-Session-Id groups related calls. There is also an async logging path if you do not want a proxy in your request path.

The header-based model is its real advantage. Attribution works from any language, any SDK, and any framework, with no vendor SDK in your application. It is the lowest-friction way to get per-user cost numbers that exists.

Costs and overheads, honestly. Helicone is Apache-2.0 and self-hostable via Docker Compose or Helm, but self-hosting means operating a real stack: a Postgres/Supabase instance for application data, ClickHouse for analytics, object storage for request bodies, plus the gateway and log-collection services. That is not a sidecar; it is a system with an on-call rotation. On the hosted side, as of August 2026 the free Hobby tier advertises 10,000 requests, 1 GB storage, 7-day retention and a 10 logs/minute ingestion cap; Pro is $79/month with usage-based storage charges and one-month retention; Team is $799/month with three-month retention and compliance features; Enterprise is custom and adds on-premises deployment and unlimited retention. Verify all of these before budgeting. An EU-hosted region exists for data-residency needs.

The proxy tradeoff is the one to think hardest about: putting a third party inline makes their availability your availability. Mitigate with the async logging mode or a self-hosted gateway, and test what your application does when the observability layer is down.

Portkey

Portkey is a gateway first and an observability product second, which makes it the natural comparison to a self-run LiteLLM proxy (covered in our existing LiteLLM articles). Metadata attaches through an x-portkey-metadata header carrying a JSON object, or through the metadata parameter in its SDKs. All values must be strings of at most 128 characters. The _user key is special: it powers user-level analytics, and if you pass a user field in an OpenAI-format request body Portkey copies it into _user automatically, with an explicit _user winning if both are present.

The open-source split is the thing to understand before you commit. The gateway itself is MIT-licensed and genuinely runnable on your own infrastructure — npx @portkey-ai/gateway, Docker, Kubernetes, Cloudflare Workers — and gives you routing, retries, fallbacks, load balancing and guardrails with a basic dashboard. The control plane that provides the full observability product and, importantly, budget enforcement is the commercial offering. As of August 2026 budget limits on API keys are documented as available to Enterprise and selected Pro customers, not to everyone on a paid plan. That is a material caveat if budgets are why you are evaluating Portkey.

As of August 2026 the free Developer tier is documented at 10,000 recorded logs per month with 3-day log retention and 30-day metric retention; Production is $49/month for 100,000 logs with $9 per additional 100,000 and 30-day log retention; Enterprise is custom. Verify current figures.

Choosing between SDK-based and proxy-based

SDK / callback instrumentation Proxy / gateway instrumentation
Coverage Only code you instrumented Everything routed through it, including code you did not write
Polyglot cost One integration per language One integration, all languages
Failure mode Degrades to no telemetry Can degrade to no service
Latency added Negligible (async export) One network hop, plus the proxy's own processing
Enforcement Advisory only Can actually block a call
Rich app context Easy — you are inside the app Only what fits in headers

The honest answer for most platform teams is both: a proxy for coverage and enforcement, SDK-level spans for the application context a header can never carry.

Layer 3: gateway metering

A gateway gives you one place where every call is counted and, crucially, one place where a call can be refused. If you already run LiteLLM, our existing LiteLLM proxy article covers virtual keys, per-key budgets and spend tracking in depth and there is no point repeating it here.

What matters for the landscape decision is the shape of the tradeoff, not the product. A gateway centralises metering and enforcement, and in exchange it becomes a hard dependency on the critical path, a component you have to scale and patch, and a place where a misconfiguration takes down every model call at once rather than one feature. It also only sees traffic that goes through it — a single team calling the provider SDK directly puts a hole in your ledger that no dashboard will flag, because absent traffic is invisible. Enforce gateway use at the network or key-issuance level, not by convention.

Layer 4: OpenTelemetry GenAI semantic conventions

If you want instrumentation that outlives your current vendor, this is it. The GenAI semantic conventions now live in their own OpenTelemetry repository and define standard span attributes, metrics and events for model calls, agents and tools.

The metric that matters for cost is gen_ai.client.token.usage — a histogram in units of {token}, with required attributes gen_ai.operation.name, gen_ai.provider.name and gen_ai.token.type (values input and output), plus gen_ai.request.model and gen_ai.response.model. Agent-shaped workloads get gen_ai.invoke_agent.inference_calls and gen_ai.invoke_agent.tool_calls, which are exactly the counters you want when you are hunting runaway loops.

On spans, token usage is carried by gen_ai.usage.input_tokens and gen_ai.usage.output_tokens, with gen_ai.usage.cache_read.input_tokens and gen_ai.usage.cache_write.input_tokens when caching is in play, and gen_ai.usage.reasoning.output_tokens for reasoning output. Identity attributes such as gen_ai.conversation.id, gen_ai.agent.name and gen_ai.tool.name let you group a whole agent run.

There is no cost or price attribute anywhere in the GenAI semantic conventions. Tokens are the standardised unit; converting them to money is your responsibility, using your own price book and your own attribute name. Plan for that, and namespace your cost attribute under your own prefix so a future spec addition does not collide with it.

The second caveat: as of August 2026 the entire GenAI convention set is marked Development, not Stable. Attribute names have already changed once (gen_ai.system became gen_ai.provider.name). Budget for a rename pass, and keep the mapping from provider response fields to attribute names in one function rather than scattered across your codebase.

from opentelemetry import trace, metrics
import anthropic
 
tracer = trace.get_tracer("myapp.llm")
meter = metrics.get_meter("myapp.llm")
 
token_usage = meter.create_histogram(
    name="gen_ai.client.token.usage",
    unit="{token}",
    description="Number of input and output tokens used",
)
 
client = anthropic.Anthropic()
MODEL = "claude-sonnet-5"
 
with tracer.start_as_current_span(f"chat {MODEL}") as span:
    span.set_attribute("gen_ai.operation.name", "chat")
    span.set_attribute("gen_ai.provider.name", "anthropic")
    span.set_attribute("gen_ai.request.model", MODEL)
 
    resp = client.messages.create(
        model=MODEL,
        max_tokens=1024,
        messages=[{"role": "user", "content": "Summarise this ticket."}],
    )
 
    u = resp.usage
    cache_read = getattr(u, "cache_read_input_tokens", 0) or 0
    cache_write = getattr(u, "cache_creation_input_tokens", 0) or 0
 
    span.set_attribute("gen_ai.response.model", resp.model)
    span.set_attribute("gen_ai.usage.input_tokens", u.input_tokens)
    span.set_attribute("gen_ai.usage.output_tokens", u.output_tokens)
    span.set_attribute("gen_ai.usage.cache_read.input_tokens", cache_read)
    span.set_attribute("gen_ai.usage.cache_write.input_tokens", cache_write)
 
    # Not part of the spec - your own namespace, your own price book.
    span.set_attribute("myapp.llm.cost_usd", price_request(resp))
 
    base = {
        "gen_ai.operation.name": "chat",
        "gen_ai.provider.name": "anthropic",
        "gen_ai.request.model": MODEL,
        "gen_ai.response.model": resp.model,
    }
    token_usage.record(u.input_tokens, {**base, "gen_ai.token.type": "input"})
    token_usage.record(u.output_tokens, {**base, "gen_ai.token.type": "output"})
 

What you give up by going OTel-native is the product. There is no prompt browser, no eval harness, no side-by-side diff, no playground. You get numbers in the observability stack you already run, joined to your HTTP and database spans, and you build the rest. For a platform team that already operates Grafana or an equivalent, that is often the right trade. For a four-person team shipping features, it is not.

The decision table

Question Choose this if... Choose that if... The tradeoff you are accepting
Self-host or SaaS? Prompts contain regulated data, or you already run ClickHouse-class infrastructure Your team is small and the data is not regulated Self-hosting trades a subscription for an on-call rotation and a storage bill that grows with traffic
Proxy or SDK? Polyglot codebase, or you need to block calls Single language, deep application context matters The proxy is on the critical path; the SDK misses anything not instrumented
Vendor schema or OTel? You expect to change vendors, or want one pane of glass with the rest of the stack You want evals, prompt management and a UI today OTel GenAI is still Development-stage and has no cost attribute
Sample or log everything? Payload volume is the cost driver You need every trace for compliance Sampled payloads mean the one bad request may not be there when you look
Where do budgets live? You already run a gateway You need per-feature logic a gateway cannot express Gateway budgets are coarse; app-level budgets are code you maintain

The sampling rule that matters

Sample payloads, never counters. Token counts and costs must be recorded for 100% of calls or your totals are wrong and your alerting is noise. Prompt and response bodies are the expensive part to store, and they are what you should sample — a common split is 100% of metrics, 100% of span metadata, and 1–10% of full payloads, with a rule that forces 100% capture on errors and on any request above a cost threshold.

The tradeoff is real: at a 5% payload sample, roughly one in twenty of the anomalies you investigate will have no prompt attached. Bias the sampler toward expensive and failed requests so the ones you actually open are the ones you kept.

PII and data residency

Prompts are the highest-risk log line most companies have ever written, because users put anything in them. Three decisions to make before the first integration, not after:

  • What leaves the process. Redact at the source, not at the vendor. A vendor-side redaction feature still means the raw text crossed the network.
  • Where it is stored. Both Helicone and the major platforms offer EU regions; self-hosting is the only option that guarantees the data never leaves your account.
  • How long it lives. Retention is a pricing lever on every SaaS tier, which means your compliance answer and your invoice are coupled. Decide retention from your policy, then price it, not the other way round.

A pragmatic pattern: log full metadata, token counts and cost for everything, and log payloads only for internal environments plus an opt-in sample in production. You keep complete cost attribution — which is what this track is about — without ever building a production corpus of customer text.

A minimum viable instrumentation schema

Whatever you pick, these are the fields that make every later question answerable. If your chosen tool cannot carry all of them, that is the gap you will be writing code to fill.

  • request_id and trace_id — join key to the rest of your system
  • tenant_id, user_id (pseudonymous), feature, environment, release
  • provider, request model, response model
  • input tokens, output tokens, cache read tokens, cache write tokens, reasoning tokens
  • service tier (standard vs batch) and attempt number
  • computed cost, and the price-book version used to compute it

The last one is the field everyone forgets. Prices change. If you do not record which price book produced a cost figure, you cannot explain why last quarter's dashboard no longer matches last quarter's invoice.

What to actually do

  1. Turn on the native usage and cost API for every provider today. It is a day of work and it is the only ground truth you will ever have.
  2. Pick one attribution layer — proxy or SDK — and make it mandatory. Two half-instrumented paths are worse than one complete one.
  3. Emit OpenTelemetry GenAI attributes alongside whatever vendor format you use. It costs a few lines and it is your exit option.
  4. Write the monthly reconciliation job before you write the dashboard. A dashboard nobody has reconciled is a rumour.

The next article in this track covers the layer none of these tools give you out of the box: turning that instrumentation into per-user and per-feature cost attribution that survives an audit.