Anthropic Message Batches and the OpenAI Batch API — the discount, the real SLA, what is safely deferrable, and the failure modes that eat the saving.
Batch processing is the least clever cost lever available to you and, for the right workload, the largest. There is no prompt engineering, no quality tradeoff, no measurement problem. You submit the same requests you would have sent synchronously, you wait, and you pay half. The entire question is whether your workload can tolerate the wait — and that question has a more interesting answer than it first appears.
Discounts, size limits and turnaround windows in this article were verified against Anthropic and OpenAI documentation as of August 2026. Verify current pricing and limits against the provider's own docs before committing to a design.The discount is 50%, on both directions
Anthropic's Message Batches API and OpenAI's Batch API both discount by 50%. The important detail — and the one that makes batch the strongest single lever in this whole track — is that the discount applies to input *and* output tokens. Most cost techniques attack one side of the ledger. This one halves the whole thing.
Because it is expressed as a flat percentage rather than a rate, the discount survives price changes. Whatever a model costs synchronously today, batch is half of that. That makes it worth designing around in a way that specific dollar figures are not.
On Anthropic, batch discounts also stack with prompt caching multipliers — the provider's documentation is explicit that the two combine. If your batch shares a large common prefix across thousands of requests, you can compound a 50% batch discount with a 0.1x cache read multiplier on the shared portion.
The SLA is not what the marketing says
This is where teams get burned. Both providers advertise a fast typical case and guarantee a much slower worst case, and you have to build for the guarantee.
| Anthropic Message Batches | OpenAI Batch API | |
|---|---|---|
| Typical turnaround | Most batches complete in under 1 hour | Within 24 hours, often more quickly |
| Hard deadline | 24 hours — unfinished requests expire | 24-hour completion window |
| Result availability | When all requests finish, or after 24 hours, whichever comes first | When the batch reaches `completed` |
| Batch size cap | 100,000 requests or 256 MB, whichever is hit first | 50,000 requests, 200 MB input file |
| Submission rate | Rate limited per usage tier | Up to 2,000 batches per hour |
| Result retention | 29 days after batch creation | Retrieved from the output file object |
Anthropic's documentation is direct about this: batches expire if processing does not complete within 24 hours, and under heavy demand "you may see more requests expiring." Expiry is a normal operating condition, not an exception. If your pipeline assumes every submitted request comes back, it will break in production rather than in testing.The consolation is that you are not billed for expired or cancelled requests on Anthropic. You lose the wall-clock time, not the money. But you do have to notice and resubmit, which means the job bookkeeping described later in this article is not optional.
What is safely deferrable
The test is not "does the user need this now." It is "what breaks if this result arrives 24 hours late, or does not arrive at all." Run every candidate workload through that question and the list sorts itself quickly.
| Workload | Batch? | Why |
|---|---|---|
| Backfilling classification or tagging over a historical corpus | Yes | No consumer is waiting; a re-run is cheap |
| Nightly summarisation or digest generation | Yes | Deadline is the next morning, not the next second |
| Eval suites and regression runs against a frozen dataset | Yes | Ideal fit — large, repetitive, no latency requirement |
| Synthetic data and training-set generation | Yes | Pure throughput work |
| Document enrichment in an ingest pipeline | Usually | Only if downstream consumers tolerate eventual availability |
| Content moderation sweeping an archive | Yes | Retrospective review, not a gate |
| Content moderation gating a live post | No | It is a synchronous decision in a user's path |
| Anything in a request/response path | No | The user is holding the connection |
| Agent loops and tool-calling turns | No | Each turn depends on the last; batch is one-shot |
| Streaming UX | No | Streaming is not available in batch mode |
| Anything where the input goes stale within a day | No | Live pricing, inventory, breaking news, session state |
The agent-loop row deserves emphasis because it is the one people try anyway. A batch request is a single model turn. If Claude returns `stop_reason: "tool_use"`, there is nobody to run the tool and continue — you would have to retrieve the batch, execute tools, and submit a second batch, turning a five-turn agent into five sequential 24-hour windows. Agent loops are not batchable. Their cost problem is a different article.
A hybrid pattern that works well: run the interactive path synchronously and the same work in batch overnight for anything the user did not explicitly request. Precomputed summaries, embeddings, and enrichment can all be produced at half price ahead of time, leaving the synchronous path to handle only genuine on-demand work.Anthropic: submit, poll, retrieve
Anthropic's API takes a list of requests inline. Each carries a `custom_id` — 1 to 64 characters, matching `^[a-zA-Z0-9_-]{1,64}$` — and a `params` object that is exactly a Messages API request body.
import anthropic
from anthropic.types.message_create_params import MessageCreateParamsNonStreaming
from anthropic.types.messages.batch_create_params import Request
client = anthropic.Anthropic()
def build_request(record_id: str, document: str) -> Request:
return Request(
custom_id=record_id, # your own primary key — this is how you rejoin results
params=MessageCreateParamsNonStreaming(
model="claude-haiku-4-5",
max_tokens=512, # must be at least 1; max_tokens=0 is not allowed in a batch
system=[{
"type": "text",
"text": CLASSIFICATION_INSTRUCTIONS,
# Batches routinely run longer than 5 minutes, so the 1-hour TTL
# is the right choice for a shared prefix inside a batch.
"cache_control": {"type": "ephemeral", "ttl": "1h"},
}],
messages=[{"role": "user", "content": document}],
),
)
batch = client.messages.batches.create(
requests=[build_request(r["id"], r["body"]) for r in records]
)
print(batch.id, batch.processing_status) # "in_progress"Validation of each request's `params` happens asynchronously, and validation errors only surface when the whole batch ends. Anthropic's docs recommend verifying your request shape against the synchronous Messages API first — send one request the normal way before submitting ten thousand of them. That five-second check has saved a lot of 24-hour round trips.
import time
# Poll until processing_status flips from "in_progress" to "ended".
while True:
batch = client.messages.batches.retrieve(batch.id)
if batch.processing_status == "ended":
break
print(batch.request_counts) # processing / succeeded / errored / canceled / expired
time.sleep(60)
# Stream results rather than downloading them all — batches can be very large.
outcomes = {"succeeded": 0, "errored": 0, "canceled": 0, "expired": 0}
needs_resubmit = []
for result in client.messages.batches.results(batch.id):
outcomes[result.result.type] += 1
if result.result.type == "succeeded":
message = result.result.message
text = next(b.text for b in message.content if b.type == "text")
save_result(result.custom_id, text)
elif result.result.type == "errored":
error_type = result.result.error.error.type
if error_type == "invalid_request":
# A bug in your request builder. Resubmitting unchanged will fail again.
record_permanent_failure(result.custom_id, error_type)
else:
needs_resubmit.append(result.custom_id)
elif result.result.type in ("expired", "canceled"):
# Not billed. Genuinely just needs to be run again.
needs_resubmit.append(result.custom_id)
print(outcomes, f"{len(needs_resubmit)} to resubmit")Results come back in arbitrary order. Both providers say so explicitly. Never zip results against your input list by position — always join on `custom_id`. This is the single most common batch bug and it produces silently mismatched data rather than an exception.OpenAI: upload a JSONL file, then create the batch
OpenAI's flow has one more step: you write a JSONL file, upload it with purpose `batch`, and reference the uploaded file. Each line has four fields — `custom_id`, `method`, `url`, and `body` — where `body` is the request payload for the target endpoint.
import json
from openai import OpenAI
client = OpenAI()
with open("batch_input.jsonl", "w", encoding="utf-8") as f:
for record in records:
f.write(json.dumps({
"custom_id": record["id"],
"method": "POST",
"url": "/v1/responses",
"body": {
"model": "gpt-5.6-sol",
"input": [
{"role": "developer", "content": CLASSIFICATION_INSTRUCTIONS},
{"role": "user", "content": record["body"]},
],
"max_output_tokens": 512,
},
}) + "\n")
input_file = client.files.create(file=open("batch_input.jsonl", "rb"), purpose="batch")
batch = client.batches.create(
input_file_id=input_file.id,
endpoint="/v1/responses",
completion_window="24h",
)
print(batch.id, batch.status) # "validating"Supported endpoints include `/v1/responses`, `/v1/chat/completions`, `/v1/embeddings`, `/v1/completions`, `/v1/moderations`, `/v1/images/generations`, `/v1/images/edits` and `/v1/videos`. Embeddings in particular are a strong batch candidate — high volume, no latency requirement, and a workload that is almost always a backfill.
import time
TERMINAL = {"completed", "failed", "expired", "cancelled"}
while True:
batch = client.batches.retrieve(batch.id)
if batch.status in TERMINAL:
break
# validating -> in_progress -> finalizing -> completed
time.sleep(60)
if batch.status == "completed":
body = client.files.content(batch.output_file_id).text
for line in body.splitlines():
row = json.loads(line)
save_result(row["custom_id"], row["response"]["body"])
# Failures land in a SEPARATE file. If you only read output_file_id you will
# silently lose every failed request.
if batch.error_file_id:
errors = client.files.content(batch.error_file_id).text
for line in errors.splitlines():
row = json.loads(line)
handle_failure(row["custom_id"], row["error"])OpenAI writes successes and failures to two different files. A pipeline that reads only `output_file_id` will process a 90%-successful batch as if it were 100% complete, and the missing 10% will never be noticed. Always check `error_file_id`.The bookkeeping batch forces on you
Synchronous calls are self-describing: you send a request, you get a response, the correlation is the function call. Batch severs that. A job submitted at 22:00 might return at 22:40 or at 22:00 the next day, into a different process, possibly after a deploy. Everything you previously got for free from the call stack, you now have to persist.
- **A durable custom_id mapping, written before submission.** Store the mapping from `custom_id` to your domain record before you call `create`, not after. If the submit call succeeds and your process dies before it writes, you have a batch running that you cannot interpret.
- **Batch state in your own database.** Batch ID, submission time, expected deadline, request count, and current status. Provider APIs will tell you the status but not what the batch was *for*.
- **A deadline monitor.** Something that notices a batch has been in progress for 23 hours and alerts before the expiry, not after. Expiry is silent from your application's point of view.
- **An idempotent resubmit path.** Expired and cancelled requests were not billed and need running again. Resubmitting an already-succeeded request costs real money, so the resubmit set has to be derived from persisted state, not from an in-memory list.
- **A permanent-failure bucket.** A validation error will fail identically on every retry. Separate "retry this" from "a human needs to look at this" or you will build an infinite loop that bills you each time round.
This is the real cost of batch processing, and it is worth being honest about it. The token discount is 50%. The engineering is a job queue, a state table, a deadline alarm and a reconciliation path. For a workload processing a few thousand requests a month, that engineering costs more than the tokens it saves. For one processing millions, it is trivially worth it. Do the arithmetic before you build.
What batch costs you
| Dimension | Cost |
|---|---|
| Output quality | None. The same model produces the same outputs. |
| Latency | Severe and non-negotiable — minutes at best, 24 hours guaranteed at worst. |
| Complexity | Significant. Job store, polling, reconciliation, resubmission, alerting. |
| Reliability | Lower. Expiry is routine; partial completion is the normal case. |
| Feature access | Reduced. No streaming; no interactive tool loop; some request-scheduling and speed options do not apply in batch mode. |
The latency cost is what makes this lever safe and also what makes it narrow. Nothing about the output degrades. But you cannot A/B a latency regression away, and you cannot partially defer a user-facing request. Either the workload can wait a day or it cannot.
Where this fits with the rest of your stack
Batch APIs handle model inference. They do not handle the compute around it — chunking documents, calling embeddings, writing to a vector store, running the fan-out. If your deferrable workload is a pipeline rather than a single model call, the AI Workshack guide to batch AI processing with Modal covers running that orchestration layer with parallel execution and cost control, and the two compose well: Modal runs the pipeline, provider batch APIs run the inference inside it at half price.
For attribution — knowing which batch, which prompt version and which customer generated which spend — put the gateway or tracing layer in front of your submission code rather than trying to reconstruct it from provider invoices later. The LiteLLM guide covers gateway-level cost tracking, and the LangFuse and LangSmith guides cover tracing token usage by prompt version.
A five-step adoption path
- Audit your traffic for requests where no human is waiting. Scheduled jobs, backfills, ingest enrichment, and eval runs are the usual candidates. Total their token spend — that number is your ceiling, and half of it is your prize.
- Verify one request synchronously first. Same model, same prompt, same schema. Asynchronous validation means a malformed batch costs you a day to discover.
- Persist the custom_id mapping before you submit anything, and store the batch ID with its deadline.
- Submit a small batch — a hundred requests — and exercise the full retrieval path including the error file and the resubmit branch. Do not discover your reconciliation bugs on a 50,000-request batch.
- Add prompt caching with a 1-hour TTL if your batch shares a large prefix, and confirm from the usage fields that the discounts are compounding rather than assuming they are.
As of August 2026 both providers discount batch at 50% with a 24-hour window. Both have changed limits and supported endpoints more than once. Re-check the size caps, the endpoint list and the retention windows against the provider docs before you design around them.