BatchRouter Docs

Chat completions

Send a single chat request to BatchRouter with POST /v1/chat/completions. Get the answer right away (with optional SSE streaming), or queue it as a one-item batch at the lower batch price and poll for the result.

POST /v1/chat/completions takes a standard OpenAI chat completions body. It gives you two ways to run a single request:

ModeHow you askWhat you getPrice
DirectThe standard body. Add "stream": true to streamThe answer in the response, or as server-sent eventsDirect price
Queued as batchAdd "batch": true or a batch object202 with a job ID. Poll for the answer or get a webhookBatch price

For many requests at once, use a batch: the native /v1/batches API or the OpenAI Batch API under /v1/openai/v1/.

Before you start

You need a BatchRouter API key and credits. See Authentication. The base URL for OpenAI SDKs is https://api.batchrouter.com/v1.

model must be a model id from the catalog (GET /v1/catalog/models), and it must match exactly. See Models & catalog.

Direct requests

A direct request returns the answer in the response, as it would from OpenAI.

curl https://api.batchrouter.com/v1/chat/completions \
  -H "Authorization: Bearer $BATCHROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.4-mini",
    "messages": [{"role": "user", "content": "Classify this ticket: refund not received"}],
    "max_tokens": 64
  }'

The response is a standard chat.completion object with usage.

Streaming

Set "stream": true to receive the answer as server-sent events (text/event-stream): one chat.completion.chunk per event, ending with data: [DONE]. Add "stream_options": {"include_usage": true} to get usage in the last chunk.

curl -N https://api.batchrouter.com/v1/chat/completions \
  -H "Authorization: Bearer $BATCHROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.4-mini",
    "messages": [{"role": "user", "content": "Write a haiku about queues"}],
    "stream": true,
    "stream_options": {"include_usage": true}
  }'

Routing fields

Two BatchRouter fields are accepted in the body, next to the OpenAI fields:

FieldMeaning
provider_preferencesonly (a list of provider slugs to allow), order (providers to try first, in order) and allow_fallbacks (whether other providers may serve the request)
privacy_tierstandard, confidential or restricted, as for batches. See Privacy tiers

OpenAI SDKs send them through extra_body (Python) or as extra fields in the request (Node):

const completion = await client.chat.completions.create({
  model: "gpt-5.4-mini",
  messages: [{ role: "user", content: "Classify this ticket: refund not received" }],
  // @ts-expect-error BatchRouter routing fields
  provider_preferences: { order: ["provider-slug"], allow_fallbacks: true },
  privacy_tier: "standard",
});

Provider slugs come from GET /v1/providers. See Providers.

How a direct request is routed and billed

  • Routing. BatchRouter sends the request to a provider lane that serves the exact model, ranked by direct price and reliability, within your privacy_tier and provider_preferences. If that lane gives a connection error, times out or returns a server error before the first byte, BatchRouter tries the next lane. Once streaming has started there is no failover.
  • Hold. Before the call, BatchRouter holds the most the request can cost: the estimated prompt tokens at the input price, plus max_tokens (or the lane's output limit) at the output price. If your balance can't cover the hold, the answer is 402 with code insufficient_funds. Setting max_tokens keeps the hold small.
  • Settle. After the call, BatchRouter charges the actual usage and releases the rest of the hold.

Direct requests cost more than batch work. If you don't need the answer right away, queue it as a batch.

Queue a request as a batch

Add "batch": true and the same request becomes a one-item batch, at the batch price, with no quote step. You get a job back straight away and collect the answer later.

{
  "model": "gpt-5.4-mini",
  "messages": [{"role": "user", "content": "Classify this ticket: refund not received"}],
  "batch": true
}

For more control, send a batch object instead of true:

FieldMeaning
completion_window"24h"
webhook{"url": "https://…", "secret": "…"}. BatchRouter calls it when the job finishes
metadataYour own key-value pairs

"stream": true can't be combined with batch. That request fails with 400 stream_not_supported_for_batch.

  1. Create the job.

    curl https://api.batchrouter.com/v1/chat/completions \
      -H "Authorization: Bearer $BATCHROUTER_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "gpt-5.4-mini",
        "messages": [{"role": "user", "content": "Classify this ticket: refund not received"}],
        "batch": {
          "completion_window": "24h",
          "webhook": {"url": "https://example.com/hooks/batchrouter", "secret": "a-long-random-shared-secret"},
          "metadata": {"ticket": "T-1042"}
        }
      }'

    The response is 202 Accepted with a job:

    {
      "id": "chatjob_9f1c2b7a4e5d40c8a3b6d8e2f4a1c7d0",
      "object": "chat.completion.job",
      "status": "queued",
      "model": "gpt-5.4-mini",
      "batch_id": "batch_9f1c2b7a4e5d40c8a3b6d8e2f4a1c7d0",
      "created": 1791460800
    }

    batch_id is the one-item batch behind the job. The two ids share the same suffix — the job chatjob_<x> is the batch batch_<x>. It shows up in your dashboard and the batch API like any other batch.

  2. Poll the job.

    GET /v1/chat/completions/{id} returns the job. status is queued, in_progress, completed, failed or cancelled.

    curl https://api.batchrouter.com/v1/chat/completions/chatjob_9f1c2b7a4e5d40c8a3b6d8e2f4a1c7d0 \
      -H "Authorization: Bearer $BATCHROUTER_API_KEY"
  3. Read the answer.

    When status is completed, result holds the standard chat.completion object. When it is failed, error says why.

    {
      "id": "chatjob_9f1c2b7a4e5d40c8a3b6d8e2f4a1c7d0",
      "object": "chat.completion.job",
      "status": "completed",
      "model": "gpt-5.4-mini",
      "batch_id": "batch_9f1c2b7a4e5d40c8a3b6d8e2f4a1c7d0",
      "created": 1791460800,
      "result": {
        "object": "chat.completion",
        "model": "gpt-5.4-mini",
        "choices": [{"index": 0, "message": {"role": "assistant", "content": "billing"}, "finish_reason": "stop"}],
        "usage": {"prompt_tokens": 18, "completion_tokens": 2, "total_tokens": 20}
      }
    }

Skip the polling with a webhook

Pass batch.webhook and BatchRouter calls your URL when the job finishes. Verify the request before you trust it, as described in Verify the signature, then fetch the job with GET /v1/chat/completions/{id}.

OpenAI SDKs can send batch too (through extra_body in Python), but they read the 202 reply as a chat completion. The examples above use plain HTTP so the job's fields are easy to read.

Errors

Errors use the standard error envelope.

StatusCodeWhen
400no_direct_laneDirect requests only: no provider serves the model for direct requests with these settings. Send "batch": true to queue it at batch price instead
400stream_not_supported_for_batch"stream": true together with batch
402insufficient_fundsDirect: your balance can't cover the hold. Queued: your credits can't fund the job
403spending_limit_reachedDirect or queued: the request would exceed your spending limit
409direct_not_cancellablePOST /v1/batches/{batch_id}/cancel on the batch behind a direct request. A direct request can't be cancelled

Next steps

On this page