Chat completions
Send a single chat request to BatchRouter with POST /v1/chat/completions. Get the answer right away (with optional SSE streaming), or queue it as a one-item batch at the lower batch price and poll for the result.
POST /v1/chat/completions takes a standard OpenAI chat completions body. It gives you two ways to
run a single request:
| Mode | How you ask | What you get | Price |
|---|---|---|---|
| Direct | The standard body. Add "stream": true to stream | The answer in the response, or as server-sent events | Direct price |
| Queued as batch | Add "batch": true or a batch object | 202 with a job ID. Poll for the answer or get a webhook | Batch price |
For many requests at once, use a batch: the native /v1/batches API or the
OpenAI Batch API under /v1/openai/v1/.
Before you start
You need a BatchRouter API key and credits. See Authentication. The
base URL for OpenAI SDKs is https://api.batchrouter.com/v1.
model must be a model id from the catalog (GET /v1/catalog/models), and it must match exactly.
See Models & catalog.
Direct requests
A direct request returns the answer in the response, as it would from OpenAI.
curl https://api.batchrouter.com/v1/chat/completions \
-H "Authorization: Bearer $BATCHROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.4-mini",
"messages": [{"role": "user", "content": "Classify this ticket: refund not received"}],
"max_tokens": 64
}'The response is a standard chat.completion object with usage.
Streaming
Set "stream": true to receive the answer as server-sent events (text/event-stream): one
chat.completion.chunk per event, ending with data: [DONE]. Add
"stream_options": {"include_usage": true} to get usage in the last chunk.
curl -N https://api.batchrouter.com/v1/chat/completions \
-H "Authorization: Bearer $BATCHROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.4-mini",
"messages": [{"role": "user", "content": "Write a haiku about queues"}],
"stream": true,
"stream_options": {"include_usage": true}
}'Routing fields
Two BatchRouter fields are accepted in the body, next to the OpenAI fields:
| Field | Meaning |
|---|---|
provider_preferences | only (a list of provider slugs to allow), order (providers to try first, in order) and allow_fallbacks (whether other providers may serve the request) |
privacy_tier | standard, confidential or restricted, as for batches. See Privacy tiers |
OpenAI SDKs send them through extra_body (Python) or as extra fields in the request (Node):
const completion = await client.chat.completions.create({
model: "gpt-5.4-mini",
messages: [{ role: "user", content: "Classify this ticket: refund not received" }],
// @ts-expect-error BatchRouter routing fields
provider_preferences: { order: ["provider-slug"], allow_fallbacks: true },
privacy_tier: "standard",
});Provider slugs come from GET /v1/providers. See Providers.
How a direct request is routed and billed
- Routing. BatchRouter sends the request to a provider lane that serves the exact model, ranked by
direct price and reliability, within your
privacy_tierandprovider_preferences. If that lane gives a connection error, times out or returns a server error before the first byte, BatchRouter tries the next lane. Once streaming has started there is no failover. - Hold. Before the call, BatchRouter holds the most the request can cost: the estimated prompt
tokens at the input price, plus
max_tokens(or the lane's output limit) at the output price. If your balance can't cover the hold, the answer is402with codeinsufficient_funds. Settingmax_tokenskeeps the hold small. - Settle. After the call, BatchRouter charges the actual
usageand releases the rest of the hold.
Direct requests cost more than batch work. If you don't need the answer right away, queue it as a batch.
Queue a request as a batch
Add "batch": true and the same request becomes a one-item batch, at the batch price, with no quote
step. You get a job back straight away and collect the answer later.
{
"model": "gpt-5.4-mini",
"messages": [{"role": "user", "content": "Classify this ticket: refund not received"}],
"batch": true
}For more control, send a batch object instead of true:
| Field | Meaning |
|---|---|
completion_window | "24h" |
webhook | {"url": "https://…", "secret": "…"}. BatchRouter calls it when the job finishes |
metadata | Your own key-value pairs |
"stream": true can't be combined with batch. That request fails with
400 stream_not_supported_for_batch.
-
Create the job.
curl https://api.batchrouter.com/v1/chat/completions \ -H "Authorization: Bearer $BATCHROUTER_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-5.4-mini", "messages": [{"role": "user", "content": "Classify this ticket: refund not received"}], "batch": { "completion_window": "24h", "webhook": {"url": "https://example.com/hooks/batchrouter", "secret": "a-long-random-shared-secret"}, "metadata": {"ticket": "T-1042"} } }'The response is
202 Acceptedwith a job:{ "id": "chatjob_9f1c2b7a4e5d40c8a3b6d8e2f4a1c7d0", "object": "chat.completion.job", "status": "queued", "model": "gpt-5.4-mini", "batch_id": "batch_9f1c2b7a4e5d40c8a3b6d8e2f4a1c7d0", "created": 1791460800 }batch_idis the one-item batch behind the job. The two ids share the same suffix — the jobchatjob_<x>is the batchbatch_<x>. It shows up in your dashboard and the batch API like any other batch. -
Poll the job.
GET /v1/chat/completions/{id}returns the job.statusisqueued,in_progress,completed,failedorcancelled.curl https://api.batchrouter.com/v1/chat/completions/chatjob_9f1c2b7a4e5d40c8a3b6d8e2f4a1c7d0 \ -H "Authorization: Bearer $BATCHROUTER_API_KEY" -
Read the answer.
When
statusiscompleted,resultholds the standardchat.completionobject. When it isfailed,errorsays why.{ "id": "chatjob_9f1c2b7a4e5d40c8a3b6d8e2f4a1c7d0", "object": "chat.completion.job", "status": "completed", "model": "gpt-5.4-mini", "batch_id": "batch_9f1c2b7a4e5d40c8a3b6d8e2f4a1c7d0", "created": 1791460800, "result": { "object": "chat.completion", "model": "gpt-5.4-mini", "choices": [{"index": 0, "message": {"role": "assistant", "content": "billing"}, "finish_reason": "stop"}], "usage": {"prompt_tokens": 18, "completion_tokens": 2, "total_tokens": 20} } }
Skip the polling with a webhook
Pass batch.webhook and BatchRouter calls your URL when the job finishes. Verify the request before
you trust it, as described in Verify the signature, then
fetch the job with GET /v1/chat/completions/{id}.
OpenAI SDKs can send batch too (through extra_body in Python), but they read the 202 reply as a
chat completion. The examples above use plain HTTP so the job's fields are easy to read.
Errors
Errors use the standard error envelope.
| Status | Code | When |
|---|---|---|
400 | no_direct_lane | Direct requests only: no provider serves the model for direct requests with these settings. Send "batch": true to queue it at batch price instead |
400 | stream_not_supported_for_batch | "stream": true together with batch |
402 | insufficient_funds | Direct: your balance can't cover the hold. Queued: your credits can't fund the job |
403 | spending_limit_reached | Direct or queued: the request would exceed your spending limit |
409 | direct_not_cancellable | POST /v1/batches/{batch_id}/cancel on the batch behind a direct request. A direct request can't be cancelled |