Anthropic Message Batches and the OpenAI Batch API — the discount, the real SLA, what is safely deferrable, and the failure modes that eat the saving.

Batch processing is the least clever cost lever available to you and, for the right workload, the largest. There is no prompt engineering, no quality tradeoff, no measurement problem. You submit the same requests you would have sent synchronously, you wait, and you pay half. The entire question is whether your workload can tolerate the wait — and that question has a more interesting answer than it first appears.

Discounts, size limits and turnaround windows in this article were verified against Anthropic and OpenAI documentation as of August 2026. Verify current pricing and limits against the provider's own docs before committing to a design.

The discount is 50%, on both directions

Anthropic's Message Batches API and OpenAI's Batch API both discount by 50%. The important detail — and the one that makes batch the strongest single lever in this whole track — is that the discount applies to input *and* output tokens. Most cost techniques attack one side of the ledger. This one halves the whole thing.

Because it is expressed as a flat percentage rather than a rate, the discount survives price changes. Whatever a model costs synchronously today, batch is half of that. That makes it worth designing around in a way that specific dollar figures are not.

On Anthropic, batch discounts also stack with prompt caching multipliers — the provider's documentation is explicit that the two combine. If your batch shares a large common prefix across thousands of requests, you can compound a 50% batch discount with a 0.1x cache read multiplier on the shared portion.

The SLA is not what the marketing says

This is where teams get burned. Both providers advertise a fast typical case and guarantee a much slower worst case, and you have to build for the guarantee.

Anthropic Message Batches OpenAI Batch API
Typical turnaround Most batches complete in under 1 hour Within 24 hours, often more quickly
Hard deadline 24 hours — unfinished requests expire 24-hour completion window
Result availability When all requests finish, or after 24 hours, whichever comes first When the batch reaches `completed`
Batch size cap 100,000 requests or 256 MB, whichever is hit first 50,000 requests, 200 MB input file
Submission rate Rate limited per usage tier Up to 2,000 batches per hour
Result retention 29 days after batch creation Retrieved from the output file object
Anthropic's documentation is direct about this: batches expire if processing does not complete within 24 hours, and under heavy demand "you may see more requests expiring." Expiry is a normal operating condition, not an exception. If your pipeline assumes every submitted request comes back, it will break in production rather than in testing.

The consolation is that you are not billed for expired or cancelled requests on Anthropic. You lose the wall-clock time, not the money. But you do have to notice and resubmit, which means the job bookkeeping described later in this article is not optional.

What is safely deferrable

The test is not "does the user need this now." It is "what breaks if this result arrives 24 hours late, or does not arrive at all." Run every candidate workload through that question and the list sorts itself quickly.

Workload Batch? Why
Backfilling classification or tagging over a historical corpus Yes No consumer is waiting; a re-run is cheap
Nightly summarisation or digest generation Yes Deadline is the next morning, not the next second
Eval suites and regression runs against a frozen dataset Yes Ideal fit — large, repetitive, no latency requirement
Synthetic data and training-set generation Yes Pure throughput work
Document enrichment in an ingest pipeline Usually Only if downstream consumers tolerate eventual availability
Content moderation sweeping an archive Yes Retrospective review, not a gate
Content moderation gating a live post No It is a synchronous decision in a user's path
Anything in a request/response path No The user is holding the connection
Agent loops and tool-calling turns No Each turn depends on the last; batch is one-shot
Streaming UX No Streaming is not available in batch mode
Anything where the input goes stale within a day No Live pricing, inventory, breaking news, session state

The agent-loop row deserves emphasis because it is the one people try anyway. A batch request is a single model turn. If Claude returns `stop_reason: "tool_use"`, there is nobody to run the tool and continue — you would have to retrieve the batch, execute tools, and submit a second batch, turning a five-turn agent into five sequential 24-hour windows. Agent loops are not batchable. Their cost problem is a different article.

A hybrid pattern that works well: run the interactive path synchronously and the same work in batch overnight for anything the user did not explicitly request. Precomputed summaries, embeddings, and enrichment can all be produced at half price ahead of time, leaving the synchronous path to handle only genuine on-demand work.

Anthropic: submit, poll, retrieve

Anthropic's API takes a list of requests inline. Each carries a `custom_id` — 1 to 64 characters, matching `^[a-zA-Z0-9_-]{1,64}$` — and a `params` object that is exactly a Messages API request body.

import anthropic
from anthropic.types.message_create_params import MessageCreateParamsNonStreaming
from anthropic.types.messages.batch_create_params import Request
 
client = anthropic.Anthropic()
 
def build_request(record_id: str, document: str) -> Request:
    return Request(
        custom_id=record_id,  # your own primary key — this is how you rejoin results
        params=MessageCreateParamsNonStreaming(
            model="claude-haiku-4-5",
            max_tokens=512,  # must be at least 1; max_tokens=0 is not allowed in a batch
            system=[{
                "type": "text",
                "text": CLASSIFICATION_INSTRUCTIONS,
                # Batches routinely run longer than 5 minutes, so the 1-hour TTL
                # is the right choice for a shared prefix inside a batch.
                "cache_control": {"type": "ephemeral", "ttl": "1h"},
            }],
            messages=[{"role": "user", "content": document}],
        ),
    )
 
batch = client.messages.batches.create(
    requests=[build_request(r["id"], r["body"]) for r in records]
)
print(batch.id, batch.processing_status)  # "in_progress"

Validation of each request's `params` happens asynchronously, and validation errors only surface when the whole batch ends. Anthropic's docs recommend verifying your request shape against the synchronous Messages API first — send one request the normal way before submitting ten thousand of them. That five-second check has saved a lot of 24-hour round trips.

import time
 
# Poll until processing_status flips from "in_progress" to "ended".
while True:
    batch = client.messages.batches.retrieve(batch.id)
    if batch.processing_status == "ended":
        break
    print(batch.request_counts)  # processing / succeeded / errored / canceled / expired
    time.sleep(60)
 
# Stream results rather than downloading them all — batches can be very large.
outcomes = {"succeeded": 0, "errored": 0, "canceled": 0, "expired": 0}
needs_resubmit = []
 
for result in client.messages.batches.results(batch.id):
    outcomes[result.result.type] += 1
 
    if result.result.type == "succeeded":
        message = result.result.message
        text = next(b.text for b in message.content if b.type == "text")
        save_result(result.custom_id, text)
 
    elif result.result.type == "errored":
        error_type = result.result.error.error.type
        if error_type == "invalid_request":
            # A bug in your request builder. Resubmitting unchanged will fail again.
            record_permanent_failure(result.custom_id, error_type)
        else:
            needs_resubmit.append(result.custom_id)
 
    elif result.result.type in ("expired", "canceled"):
        # Not billed. Genuinely just needs to be run again.
        needs_resubmit.append(result.custom_id)
 
print(outcomes, f"{len(needs_resubmit)} to resubmit")
Results come back in arbitrary order. Both providers say so explicitly. Never zip results against your input list by position — always join on `custom_id`. This is the single most common batch bug and it produces silently mismatched data rather than an exception.

OpenAI: upload a JSONL file, then create the batch

OpenAI's flow has one more step: you write a JSONL file, upload it with purpose `batch`, and reference the uploaded file. Each line has four fields — `custom_id`, `method`, `url`, and `body` — where `body` is the request payload for the target endpoint.

import json
from openai import OpenAI
 
client = OpenAI()
 
with open("batch_input.jsonl", "w", encoding="utf-8") as f:
    for record in records:
        f.write(json.dumps({
            "custom_id": record["id"],
            "method": "POST",
            "url": "/v1/responses",
            "body": {
                "model": "gpt-5.6-sol",
                "input": [
                    {"role": "developer", "content": CLASSIFICATION_INSTRUCTIONS},
                    {"role": "user", "content": record["body"]},
                ],
                "max_output_tokens": 512,
            },
        }) + "\n")
 
input_file = client.files.create(file=open("batch_input.jsonl", "rb"), purpose="batch")
 
batch = client.batches.create(
    input_file_id=input_file.id,
    endpoint="/v1/responses",
    completion_window="24h",
)
print(batch.id, batch.status)  # "validating"

Supported endpoints include `/v1/responses`, `/v1/chat/completions`, `/v1/embeddings`, `/v1/completions`, `/v1/moderations`, `/v1/images/generations`, `/v1/images/edits` and `/v1/videos`. Embeddings in particular are a strong batch candidate — high volume, no latency requirement, and a workload that is almost always a backfill.

import time
 
TERMINAL = {"completed", "failed", "expired", "cancelled"}
 
while True:
    batch = client.batches.retrieve(batch.id)
    if batch.status in TERMINAL:
        break
    # validating -> in_progress -> finalizing -> completed
    time.sleep(60)
 
if batch.status == "completed":
    body = client.files.content(batch.output_file_id).text
    for line in body.splitlines():
        row = json.loads(line)
        save_result(row["custom_id"], row["response"]["body"])
 
# Failures land in a SEPARATE file. If you only read output_file_id you will
# silently lose every failed request.
if batch.error_file_id:
    errors = client.files.content(batch.error_file_id).text
    for line in errors.splitlines():
        row = json.loads(line)
        handle_failure(row["custom_id"], row["error"])
OpenAI writes successes and failures to two different files. A pipeline that reads only `output_file_id` will process a 90%-successful batch as if it were 100% complete, and the missing 10% will never be noticed. Always check `error_file_id`.

The bookkeeping batch forces on you

Synchronous calls are self-describing: you send a request, you get a response, the correlation is the function call. Batch severs that. A job submitted at 22:00 might return at 22:40 or at 22:00 the next day, into a different process, possibly after a deploy. Everything you previously got for free from the call stack, you now have to persist.

  • **A durable custom_id mapping, written before submission.** Store the mapping from `custom_id` to your domain record before you call `create`, not after. If the submit call succeeds and your process dies before it writes, you have a batch running that you cannot interpret.
  • **Batch state in your own database.** Batch ID, submission time, expected deadline, request count, and current status. Provider APIs will tell you the status but not what the batch was *for*.
  • **A deadline monitor.** Something that notices a batch has been in progress for 23 hours and alerts before the expiry, not after. Expiry is silent from your application's point of view.
  • **An idempotent resubmit path.** Expired and cancelled requests were not billed and need running again. Resubmitting an already-succeeded request costs real money, so the resubmit set has to be derived from persisted state, not from an in-memory list.
  • **A permanent-failure bucket.** A validation error will fail identically on every retry. Separate "retry this" from "a human needs to look at this" or you will build an infinite loop that bills you each time round.

This is the real cost of batch processing, and it is worth being honest about it. The token discount is 50%. The engineering is a job queue, a state table, a deadline alarm and a reconciliation path. For a workload processing a few thousand requests a month, that engineering costs more than the tokens it saves. For one processing millions, it is trivially worth it. Do the arithmetic before you build.

What batch costs you

Dimension Cost
Output quality None. The same model produces the same outputs.
Latency Severe and non-negotiable — minutes at best, 24 hours guaranteed at worst.
Complexity Significant. Job store, polling, reconciliation, resubmission, alerting.
Reliability Lower. Expiry is routine; partial completion is the normal case.
Feature access Reduced. No streaming; no interactive tool loop; some request-scheduling and speed options do not apply in batch mode.

The latency cost is what makes this lever safe and also what makes it narrow. Nothing about the output degrades. But you cannot A/B a latency regression away, and you cannot partially defer a user-facing request. Either the workload can wait a day or it cannot.

Where this fits with the rest of your stack

Batch APIs handle model inference. They do not handle the compute around it — chunking documents, calling embeddings, writing to a vector store, running the fan-out. If your deferrable workload is a pipeline rather than a single model call, the AI Workshack guide to batch AI processing with Modal covers running that orchestration layer with parallel execution and cost control, and the two compose well: Modal runs the pipeline, provider batch APIs run the inference inside it at half price.

For attribution — knowing which batch, which prompt version and which customer generated which spend — put the gateway or tracing layer in front of your submission code rather than trying to reconstruct it from provider invoices later. The LiteLLM guide covers gateway-level cost tracking, and the LangFuse and LangSmith guides cover tracing token usage by prompt version.

A five-step adoption path

  1. Audit your traffic for requests where no human is waiting. Scheduled jobs, backfills, ingest enrichment, and eval runs are the usual candidates. Total their token spend — that number is your ceiling, and half of it is your prize.
  2. Verify one request synchronously first. Same model, same prompt, same schema. Asynchronous validation means a malformed batch costs you a day to discover.
  3. Persist the custom_id mapping before you submit anything, and store the batch ID with its deadline.
  4. Submit a small batch — a hundred requests — and exercise the full retrieval path including the error file and the resubmit branch. Do not discover your reconciliation bugs on a 50,000-request batch.
  5. Add prompt caching with a 1-hour TTL if your batch shares a large prefix, and confirm from the usage fields that the discounts are compounding rather than assuming they are.
As of August 2026 both providers discount batch at 50% with a 24-hour window. Both have changed limits and supported endpoints more than once. Re-check the size caps, the endpoint list and the retention windows against the provider docs before you design around them.