BatchRouter Docs

Provider API

The API a BatchRouter provider hosts. Batch work arrives through the OpenAI Batch API with BatchRouter extensions for deadlines, pay, acceptance and partial results. Direct customer requests arrive through chat completions.

This page is the full specification of the API your server hosts to receive work from BatchRouter. There are two kinds of work, and each has its own endpoint:

  • Batch work arrives through the OpenAI Batch API, hosted by you, with a few BatchRouter extensions. You get the whole batch and its deadline, so you can hold the jobs and run them in spare capacity. This is where most of the volume is.
  • Direct requests arrive through POST /v1/chat/completions. BatchRouter calls it only when a customer sends a single request and wants the answer right away, never for batch work.

Which of the two a lane receives depends on its integration:

Integrationintegration at registration, api_protocol on the laneBatch workDirect requests
OpenAI Batch API (recommended)openai_batchYesYes, at the lane's direct price
Chat completions onlyopenai_compatNoYes, at the lane's price

A chat-only lane gets no batch work. To receive batch work, host the batch endpoints below and switch the lane to openai_batch.

Integrations built on BatchRouter's earlier batch contract keep working and will be migrated to this API. New providers can't choose that contract.

Endpoints

Host these endpoints at your base URL. The base URL is the https:// URL you enter in the portal, without a trailing /v1.

EndpointStandard or extensionWhat BatchRouter uses it for
GET /v1/modelsStandardGo-live check and daily check. Must list the Model ID of every model on the lane
GET /v1/batches?limit=1StandardGo-live check that the batch API exists. Creates no work
POST /v1/filesStandardUpload the batch input (purpose=batch, multipart, JSONL)
POST /v1/batchesStandard, with extension keys in metadataCreate the batch. The response is your acceptance or decline
GET /v1/batches/{batch_id}StandardStatus and request_counts
GET /v1/batches/{batch_id}/resultsExtension: partial resultsFinished lines before the batch ends
GET /v1/files/{file_id}/contentStandardOutput and error files at the end
POST /v1/batches/{batch_id}/cancelStandardCancel
POST /v1/chat/completionsStandardDirect customer requests only

A chat-only (openai_compat) lane needs GET /v1/models and POST /v1/chat/completions only.

Authentication

BatchRouter sends Authorization: Bearer <your API key> on every call. It's the key you store in the provider portal, where it's encrypted. Use a dedicated key you can revoke without affecting your other customers.

Model IDs

The model BatchRouter sends is your Model ID for that model, exactly as you entered it in the portal and exactly as GET /v1/models lists it. See Add your models.

Batch work

A batch takes four steps: BatchRouter uploads the input file, creates the batch, polls it and reads the results. The shapes are the OpenAI Batch API's, so a server that already speaks it needs only the extensions on this page.

Input file

BatchRouter uploads one JSONL file per batch with POST /v1/files (purpose=batch, multipart). Each line is one request. All lines in a file are for one model and one endpoint.

{"custom_id":"item_7d2f8a1c4b9e40d3a6c5e8f1b2d4a090","method":"POST","url":"/v1/chat/completions","body":{"model":"llama-3.1-8b-instruct","messages":[{"role":"user","content":"Classify this ticket: refund not received"}],"max_tokens":512}}
  • url is /v1/chat/completions or /v1/embeddings.
  • body is a standard request body for that endpoint. It never contains stream.
  • custom_id is BatchRouter's item ID. It is opaque, at most 64 characters and unique within the batch. Echo it exactly in the result for that line — a result whose custom_id is anything else is ignored and flags the lane. See Identity rules.

Create the batch

BatchRouter creates the batch with POST /v1/batches:

POST /v1/batches
Authorization: Bearer <your API key>
Content-Type: application/json
Idempotency-Key: wu_3e6c9b2d5a8f41e0b7d4c1a9f0e2b385:9f2c4e6a8b1d3f50

{
  "input_file_id": "file-abc123",
  "endpoint": "/v1/chat/completions",
  "completion_window": "24h",
  "metadata": {
    "batchrouter_work_unit_id": "wu_3e6c9b2d5a8f41e0b7d4c1a9f0e2b385",
    "batchrouter_deadline_at": "2026-10-08T14:30:00Z",
    "batchrouter_window_seconds": "7200",
    "batchrouter_tier": "standard",
    "batchrouter_role": "standard",
    "batchrouter_pay_input_per_1m": "0.0400",
    "batchrouter_pay_output_per_1m": "0.0800",
    "batchrouter_pay_cached_input_per_1m": "0.0200"
  }
}

completion_window is always "24h", for compatibility. The real deadline is in metadata.

Idempotency-Key. The key is the work unit ID plus a short hash of the exact set of items being sent, so a retry of the same create (for example after a network timeout) carries the same key, while a later dispatch of a different item subset gets a new one. Treat it as opaque. If you receive a key you have already seen, return the batch you already created for it. Never create a second batch for the same key.

Metadata keys. The extensions are strings in metadata, so a stock OpenAI-compatible server accepts them unchanged.

KeyMeaning
batchrouter_work_unit_idBatchRouter's ID for this piece of work
batchrouter_deadline_atRFC 3339 time by which every line must be finished. Lines finished later are not paid
batchrouter_window_secondsYour processing window for this assignment, in seconds
batchrouter_tierThe customer tier. standard today
batchrouter_roleWhy this lane got the work. standard today. backstop is reserved for rescue work — reassigned items BatchRouter needs finished quickly. Rescue work may also go to lanes with a direct price
batchrouter_pay_input_per_1mWhat BatchRouter pays for this batch's input tokens, in USD per 1 million tokens
batchrouter_pay_output_per_1mThe same for output tokens
batchrouter_pay_cached_input_per_1mThe same for cached input tokens

Accept or decline

Your response to the create request is your answer. There is no separate quote or accept step.

Your responseWhat it means
2xx with a batch objectAccepted. The window counts from when BatchRouter sent the request
429 or 503, optionally with Retry-AfterDeclined for capacity. BatchRouter routes the batch to another lane. It counts as a decline, not a failure
Any other 4xxRejected. BatchRouter flags the lane
No response within 10 secondsTreated as a decline

Decline when you can't finish the batch by batchrouter_deadline_at. Lines that finish after the deadline are not paid, and their items are reassigned.

Status

BatchRouter polls GET /v1/batches/{batch_id} and reads status and request_counts.

Your statusWhat BatchRouter does
validating, in_progress, finalizingKeeps polling. Reads partial results if the lane supports them
cancellingKeeps polling until the batch reaches a terminal status
completedReads output_file_id and error_file_id and applies every line
failed, expired, cancelledReads whatever output and error files exist, then fails every item that has no result

Results

Each finished line has the OpenAI output-file shape, keyed by custom_id:

{"id":"batch_req_7f3a91","custom_id":"item_7d2f8a1c4b9e40d3a6c5e8f1b2d4a090","response":{"status_code":200,"request_id":"req_5c2e","body":{"id":"chatcmpl-91b","object":"chat.completion","created":1791460800,"model":"llama-3.1-8b-instruct","choices":[{"index":0,"message":{"role":"assistant","content":"billing"},"finish_reason":"stop"}],"usage":{"prompt_tokens":18,"completion_tokens":2,"total_tokens":20}}},"error":null}
  • A 2xx status_code with a null error is a success.
  • A non-2xx status_code or a non-null error fails that item. BatchRouter may retry it on another lane.
  • usage in body is required. BatchRouter pays on it. If it is missing, BatchRouter estimates the tokens and flags the lane.

Put successful lines in the output file and failed lines in the error file, as the OpenAI Batch API does. BatchRouter reads at most 20 MB per result file. Lines past the cap are never applied, so keep each file under it.

Partial results

Partial results are a BatchRouter extension. They let you hand over lines as soon as they finish, so customers see those items done before the whole batch ends.

GET /v1/batches/{batch_id}/results?after={line_id}&limit={n}
Authorization: Bearer <your API key>
ParameterMeaning
afterThe id of the last line BatchRouter has received. Omitted on the first call
limitLines per page. Default 100, maximum 1000

The response is a list of result lines in the same shape as the output and error files:

{
  "object": "list",
  "data": [
    {"id":"batch_req_7f3a91","custom_id":"item_7d2f8a1c4b9e40d3a6c5e8f1b2d4a090","response":{"status_code":200,"request_id":"req_5c2e","body":{"object":"chat.completion","choices":[{"index":0,"message":{"role":"assistant","content":"billing"},"finish_reason":"stop"}],"usage":{"prompt_tokens":18,"completion_tokens":2,"total_tokens":20}}},"error":null},
    {"id":"batch_req_7f3a92","custom_id":"item_7d2f8a1c4b9e40d3a6c5e8f1b2d4a091","response":null,"error":{"code":"context_length_exceeded","message":"Prompt is 9,214 tokens; the limit is 8,192."}}
  ],
  "first_id": "batch_req_7f3a91",
  "last_id": "batch_req_7f3a92",
  "has_more": false
}

Rules:

  • Lines are append-only and in the order they finished. Each line has a stable id.
  • Successful and failed lines both appear.
  • Each custom_id appears at most once.
  • A line returned here must be identical to the same line in the final output or error file.
  • BatchRouter polls this endpoint while the batch is validating, in_progress or finalizing, and applies each line as it arrives.

Support is set per lane: turn on partial results (partial_results on the lane) in the portal or at registration. If the endpoint answers 404 or 405, BatchRouter stops calling it for that batch and waits for the final files.

Identity rules

  • custom_id echoes the item ID BatchRouter sent, exactly. Any other value — a truncated, re-cased or self-assigned id — is ignored and flags the lane.
  • One result per custom_id.
  • An exact duplicate of a line already received is ignored.
  • A conflicting duplicate (same custom_id, different content) keeps the first result and flags the lane.
  • When the batch reaches a terminal status (completed, failed, expired, cancelled), or once batchrouter_deadline_at has passed, every item without a result fails with provider_missing_result and can be reassigned to another lane.

Cancel

BatchRouter cancels with POST /v1/batches/{batch_id}/cancel. Lines finished before the cancel are still delivered and paid, so keep them in your results.

Usage and pay

BatchRouter pays for each line from its usage, at the rates in that batch's batchrouter_pay_* metadata:

  • Chat completions: prompt_tokens and completion_tokens. Report cached input tokens in usage.prompt_tokens_details.cached_tokens.
  • Embeddings: prompt_tokens.
  • Lines finished after batchrouter_deadline_at are not paid.

Earnings are settled monthly. See Earnings and payouts.

Direct requests

BatchRouter calls POST /v1/chat/completions only for a customer's direct request: one request whose answer the customer is waiting for. It never sends batch work there.

The body is a standard chat completions body, with your Model ID as model.

  • Customer doesn't stream: BatchRouter sends "stream": false and expects a chat.completion object with usage.
  • Customer streams: BatchRouter sends "stream": true with "stream_options": {"include_usage": true} and relays your chunks to the customer as they arrive. Answer with Content-Type: text/event-stream: one data: {chat.completion.chunk} event per chunk, a final chunk that carries usage, then data: [DONE].

Streaming is server-sent events over a normal HTTP response.

Turn streaming on for a lane

In the provider portal, open the lane and tick Streams chat completions. Through the API, set chat_streaming: true when you create or update the lane (POST and PATCH on /v1/provider-portal/offerings). It applies to any lane that serves direct requests, whether openai_compat or openai_batch.

When a customer streams, BatchRouter tries the lanes that have it on first. A lane that has not turned it on still takes streamed requests, after those, so nothing breaks if you leave it off. Turn it on once your server answers stream: true as described above. Without it, a streamed request that reaches your lane costs a failover if you answer with a plain JSON body.

What to expect:

  • If your endpoint gives a connection error, times out or returns a 5xx before the first byte, BatchRouter fails over to another lane. Once you have started streaming there is no failover, so an error mid-stream fails the customer's request.
  • usage is required, as for batch work. If it is missing, BatchRouter estimates the tokens, records them as estimated, and pays and bills on the estimate. A lane that keeps answering without usage can be put under review.

Prices

The price on an openai_batch lane is its batch price, what you earn per 1 million tokens of batch work.

Every openai_batch lane also sets a direct price — direct_input_per_1m, direct_output_per_1m and, optionally, direct_cached_input_per_1m. It is required, and it does two jobs: it is what you earn for direct requests on the lane, and it is the lane's reference price — each batch price must stay below its direct price.

On a chat-only (openai_compat) lane there are no separate direct fields: the lane's regular price is its direct price.

BatchRouter adds its platform fee on top when it prices work for customers, so the price you set is what you earn.

Go-live checks

The portal's endpoint test and Go live call your endpoint with your key. Neither creates work or costs anything.

IntegrationCheckWhat passes
openai_batchGET {base URL}/v1/modelsAn OpenAI-style model list that includes the Model ID of every model on the lane
openai_batchGET {base URL}/v1/batches?limit=1An OpenAI-style list, {"object": "list", "data": [...]}. An empty data array passes
openai_compatGET {base URL}/v1/modelsThe same model list

After you go live, BatchRouter repeats the model check every day. See Test your endpoint for the usual failures and fixes.

Capacity and heartbeats

Capacity tells BatchRouter how much work a lane can take. You can set it in the portal, or from your own systems through the API:

EndpointWhat it does
GET https://api.batchrouter.com/v1/provider-portal/offeringsLists your lanes and their IDs
POST https://api.batchrouter.com/v1/provider-portal/offerings/{offering_id}/capacitySets the lane's standing capacity
POST https://api.batchrouter.com/v1/provider-portal/offerings/{offering_id}/heartbeatUpdates capacity and keeps it open for a time-to-live

Authenticate with a BatchRouter API key from the account that owns the provider, and name the provider in a header:

curl -X POST "https://api.batchrouter.com/v1/provider-portal/offerings/$OFFERING_ID/heartbeat" \
  -H "Authorization: Bearer $BATCHROUTER_API_KEY" \
  -H "x-batchrouter-provider-id: your-provider-slug" \
  -H "Content-Type: application/json" \
  -d '{"available_queue_items": 5000, "heartbeat_ttl_seconds": 300}'

Both endpoints take the capacity fields from Publish capacity, such as max_batch_items, max_concurrent_batches, available_queue_items and available_queue_tokens. A heartbeat also takes heartbeat_ttl_seconds, from 30 to 3600. Without it, the lane's configured time-to-live applies.

Once you send heartbeats, keep sending them

Each heartbeat sets an expiry: the lane's capacity is open until the time-to-live runs out. If the heartbeats stop, the lane stops taking work. To go back to standing capacity, call /capacity with "capacity_expires_at": null and "last_heartbeat_at": null.

Running on vLLM

vLLM serves GET /v1/models and POST /v1/chat/completions, so it covers direct requests and the model check as it is. For batch work you run a batch front end in front of it that hosts the batch endpoints on this page. vllm run-batch reads the same JSONL input format and writes the same output line format, which makes it a possible starting point for that front end.

Next steps

On this page