Provider API
The API a BatchRouter provider hosts. Batch work arrives through the OpenAI Batch API with BatchRouter extensions for deadlines, pay, acceptance and partial results. Direct customer requests arrive through chat completions.
This page is the full specification of the API your server hosts to receive work from BatchRouter. There are two kinds of work, and each has its own endpoint:
- Batch work arrives through the OpenAI Batch API, hosted by you, with a few BatchRouter extensions. You get the whole batch and its deadline, so you can hold the jobs and run them in spare capacity. This is where most of the volume is.
- Direct requests arrive through
POST /v1/chat/completions. BatchRouter calls it only when a customer sends a single request and wants the answer right away, never for batch work.
Which of the two a lane receives depends on its integration:
| Integration | integration at registration, api_protocol on the lane | Batch work | Direct requests |
|---|---|---|---|
| OpenAI Batch API (recommended) | openai_batch | Yes | Yes, at the lane's direct price |
| Chat completions only | openai_compat | No | Yes, at the lane's price |
A chat-only lane gets no batch work. To receive batch work, host the batch endpoints below and switch
the lane to openai_batch.
Integrations built on BatchRouter's earlier batch contract keep working and will be migrated to this API. New providers can't choose that contract.
Endpoints
Host these endpoints at your base URL. The base URL is the https:// URL you enter in the portal,
without a trailing /v1.
| Endpoint | Standard or extension | What BatchRouter uses it for |
|---|---|---|
GET /v1/models | Standard | Go-live check and daily check. Must list the Model ID of every model on the lane |
GET /v1/batches?limit=1 | Standard | Go-live check that the batch API exists. Creates no work |
POST /v1/files | Standard | Upload the batch input (purpose=batch, multipart, JSONL) |
POST /v1/batches | Standard, with extension keys in metadata | Create the batch. The response is your acceptance or decline |
GET /v1/batches/{batch_id} | Standard | Status and request_counts |
GET /v1/batches/{batch_id}/results | Extension: partial results | Finished lines before the batch ends |
GET /v1/files/{file_id}/content | Standard | Output and error files at the end |
POST /v1/batches/{batch_id}/cancel | Standard | Cancel |
POST /v1/chat/completions | Standard | Direct customer requests only |
A chat-only (openai_compat) lane needs GET /v1/models and POST /v1/chat/completions only.
Authentication
BatchRouter sends Authorization: Bearer <your API key> on every call. It's the key you store in the
provider portal, where it's encrypted. Use a dedicated key
you can revoke without affecting your other customers.
Model IDs
The model BatchRouter sends is your Model ID for that model, exactly as you entered it in the
portal and exactly as GET /v1/models lists it. See
Add your models.
Batch work
A batch takes four steps: BatchRouter uploads the input file, creates the batch, polls it and reads the results. The shapes are the OpenAI Batch API's, so a server that already speaks it needs only the extensions on this page.
Input file
BatchRouter uploads one JSONL file per batch with POST /v1/files (purpose=batch, multipart). Each
line is one request. All lines in a file are for one model and one endpoint.
{"custom_id":"item_7d2f8a1c4b9e40d3a6c5e8f1b2d4a090","method":"POST","url":"/v1/chat/completions","body":{"model":"llama-3.1-8b-instruct","messages":[{"role":"user","content":"Classify this ticket: refund not received"}],"max_tokens":512}}urlis/v1/chat/completionsor/v1/embeddings.bodyis a standard request body for that endpoint. It never containsstream.custom_idis BatchRouter's item ID. It is opaque, at most 64 characters and unique within the batch. Echo it exactly in the result for that line — a result whosecustom_idis anything else is ignored and flags the lane. See Identity rules.
Create the batch
BatchRouter creates the batch with POST /v1/batches:
POST /v1/batches
Authorization: Bearer <your API key>
Content-Type: application/json
Idempotency-Key: wu_3e6c9b2d5a8f41e0b7d4c1a9f0e2b385:9f2c4e6a8b1d3f50
{
"input_file_id": "file-abc123",
"endpoint": "/v1/chat/completions",
"completion_window": "24h",
"metadata": {
"batchrouter_work_unit_id": "wu_3e6c9b2d5a8f41e0b7d4c1a9f0e2b385",
"batchrouter_deadline_at": "2026-10-08T14:30:00Z",
"batchrouter_window_seconds": "7200",
"batchrouter_tier": "standard",
"batchrouter_role": "standard",
"batchrouter_pay_input_per_1m": "0.0400",
"batchrouter_pay_output_per_1m": "0.0800",
"batchrouter_pay_cached_input_per_1m": "0.0200"
}
}completion_window is always "24h", for compatibility. The real deadline is in metadata.
Idempotency-Key. The key is the work unit ID plus a short hash of the exact set of items being
sent, so a retry of the same create (for example after a network timeout) carries the same key, while
a later dispatch of a different item subset gets a new one. Treat it as opaque. If you receive a key
you have already seen, return the batch you already created for it. Never create a second batch for
the same key.
Metadata keys. The extensions are strings in metadata, so a stock OpenAI-compatible server
accepts them unchanged.
| Key | Meaning |
|---|---|
batchrouter_work_unit_id | BatchRouter's ID for this piece of work |
batchrouter_deadline_at | RFC 3339 time by which every line must be finished. Lines finished later are not paid |
batchrouter_window_seconds | Your processing window for this assignment, in seconds |
batchrouter_tier | The customer tier. standard today |
batchrouter_role | Why this lane got the work. standard today. backstop is reserved for rescue work — reassigned items BatchRouter needs finished quickly. Rescue work may also go to lanes with a direct price |
batchrouter_pay_input_per_1m | What BatchRouter pays for this batch's input tokens, in USD per 1 million tokens |
batchrouter_pay_output_per_1m | The same for output tokens |
batchrouter_pay_cached_input_per_1m | The same for cached input tokens |
Accept or decline
Your response to the create request is your answer. There is no separate quote or accept step.
| Your response | What it means |
|---|---|
2xx with a batch object | Accepted. The window counts from when BatchRouter sent the request |
429 or 503, optionally with Retry-After | Declined for capacity. BatchRouter routes the batch to another lane. It counts as a decline, not a failure |
Any other 4xx | Rejected. BatchRouter flags the lane |
| No response within 10 seconds | Treated as a decline |
Decline when you can't finish the batch by batchrouter_deadline_at. Lines that finish after the
deadline are not paid, and their items are reassigned.
Status
BatchRouter polls GET /v1/batches/{batch_id} and reads status and request_counts.
Your status | What BatchRouter does |
|---|---|
validating, in_progress, finalizing | Keeps polling. Reads partial results if the lane supports them |
cancelling | Keeps polling until the batch reaches a terminal status |
completed | Reads output_file_id and error_file_id and applies every line |
failed, expired, cancelled | Reads whatever output and error files exist, then fails every item that has no result |
Results
Each finished line has the OpenAI output-file shape, keyed by custom_id:
{"id":"batch_req_7f3a91","custom_id":"item_7d2f8a1c4b9e40d3a6c5e8f1b2d4a090","response":{"status_code":200,"request_id":"req_5c2e","body":{"id":"chatcmpl-91b","object":"chat.completion","created":1791460800,"model":"llama-3.1-8b-instruct","choices":[{"index":0,"message":{"role":"assistant","content":"billing"},"finish_reason":"stop"}],"usage":{"prompt_tokens":18,"completion_tokens":2,"total_tokens":20}}},"error":null}- A
2xxstatus_codewith anullerroris a success. - A non-2xx
status_codeor a non-nullerrorfails that item. BatchRouter may retry it on another lane. usageinbodyis required. BatchRouter pays on it. If it is missing, BatchRouter estimates the tokens and flags the lane.
Put successful lines in the output file and failed lines in the error file, as the OpenAI Batch API does. BatchRouter reads at most 20 MB per result file. Lines past the cap are never applied, so keep each file under it.
Partial results
Partial results are a BatchRouter extension. They let you hand over lines as soon as they finish, so customers see those items done before the whole batch ends.
GET /v1/batches/{batch_id}/results?after={line_id}&limit={n}
Authorization: Bearer <your API key>| Parameter | Meaning |
|---|---|
after | The id of the last line BatchRouter has received. Omitted on the first call |
limit | Lines per page. Default 100, maximum 1000 |
The response is a list of result lines in the same shape as the output and error files:
{
"object": "list",
"data": [
{"id":"batch_req_7f3a91","custom_id":"item_7d2f8a1c4b9e40d3a6c5e8f1b2d4a090","response":{"status_code":200,"request_id":"req_5c2e","body":{"object":"chat.completion","choices":[{"index":0,"message":{"role":"assistant","content":"billing"},"finish_reason":"stop"}],"usage":{"prompt_tokens":18,"completion_tokens":2,"total_tokens":20}}},"error":null},
{"id":"batch_req_7f3a92","custom_id":"item_7d2f8a1c4b9e40d3a6c5e8f1b2d4a091","response":null,"error":{"code":"context_length_exceeded","message":"Prompt is 9,214 tokens; the limit is 8,192."}}
],
"first_id": "batch_req_7f3a91",
"last_id": "batch_req_7f3a92",
"has_more": false
}Rules:
- Lines are append-only and in the order they finished. Each line has a stable
id. - Successful and failed lines both appear.
- Each
custom_idappears at most once. - A line returned here must be identical to the same line in the final output or error file.
- BatchRouter polls this endpoint while the batch is
validating,in_progressorfinalizing, and applies each line as it arrives.
Support is set per lane: turn on partial results (partial_results on the lane) in the portal or
at registration. If the endpoint answers 404 or 405, BatchRouter stops calling it for that batch
and waits for the final files.
Identity rules
custom_idechoes the item ID BatchRouter sent, exactly. Any other value — a truncated, re-cased or self-assigned id — is ignored and flags the lane.- One result per
custom_id. - An exact duplicate of a line already received is ignored.
- A conflicting duplicate (same
custom_id, different content) keeps the first result and flags the lane. - When the batch reaches a terminal status (
completed,failed,expired,cancelled), or oncebatchrouter_deadline_athas passed, every item without a result fails withprovider_missing_resultand can be reassigned to another lane.
Cancel
BatchRouter cancels with POST /v1/batches/{batch_id}/cancel. Lines finished before the cancel are
still delivered and paid, so keep them in your results.
Usage and pay
BatchRouter pays for each line from its usage, at the rates in that batch's batchrouter_pay_*
metadata:
- Chat completions:
prompt_tokensandcompletion_tokens. Report cached input tokens inusage.prompt_tokens_details.cached_tokens. - Embeddings:
prompt_tokens. - Lines finished after
batchrouter_deadline_atare not paid.
Earnings are settled monthly. See Earnings and payouts.
Direct requests
BatchRouter calls POST /v1/chat/completions only for a customer's direct request: one request
whose answer the customer is waiting for. It never sends batch work there.
The body is a standard chat completions body, with your Model ID as model.
- Customer doesn't stream: BatchRouter sends
"stream": falseand expects achat.completionobject withusage. - Customer streams: BatchRouter sends
"stream": truewith"stream_options": {"include_usage": true}and relays your chunks to the customer as they arrive. Answer withContent-Type: text/event-stream: onedata: {chat.completion.chunk}event per chunk, a final chunk that carriesusage, thendata: [DONE].
Streaming is server-sent events over a normal HTTP response.
Turn streaming on for a lane
In the provider portal, open the lane and tick Streams chat completions. Through the API, set
chat_streaming: true when you create or update the lane (POST and PATCH on
/v1/provider-portal/offerings). It applies to any lane that serves direct requests, whether
openai_compat or openai_batch.
When a customer streams, BatchRouter tries the lanes that have it on first. A lane that has not
turned it on still takes streamed requests, after those, so nothing breaks if you leave it off. Turn
it on once your server answers stream: true as described above. Without it, a streamed request that
reaches your lane costs a failover if you answer with a plain JSON body.
What to expect:
- If your endpoint gives a connection error, times out or returns a
5xxbefore the first byte, BatchRouter fails over to another lane. Once you have started streaming there is no failover, so an error mid-stream fails the customer's request. usageis required, as for batch work. If it is missing, BatchRouter estimates the tokens, records them as estimated, and pays and bills on the estimate. A lane that keeps answering withoutusagecan be put under review.
Prices
The price on an openai_batch lane is its batch price, what you earn per 1 million tokens of batch
work.
Every openai_batch lane also sets a direct price — direct_input_per_1m,
direct_output_per_1m and, optionally, direct_cached_input_per_1m. It is required, and it
does two jobs: it is what you earn for direct requests on the lane, and it is the lane's reference
price — each batch price must stay below its direct price.
On a chat-only (openai_compat) lane there are no separate direct fields: the lane's regular price
is its direct price.
BatchRouter adds its platform fee on top when it prices work for customers, so the price you set is what you earn.
Go-live checks
The portal's endpoint test and Go live call your endpoint with your key. Neither creates work or costs anything.
| Integration | Check | What passes |
|---|---|---|
openai_batch | GET {base URL}/v1/models | An OpenAI-style model list that includes the Model ID of every model on the lane |
openai_batch | GET {base URL}/v1/batches?limit=1 | An OpenAI-style list, {"object": "list", "data": [...]}. An empty data array passes |
openai_compat | GET {base URL}/v1/models | The same model list |
After you go live, BatchRouter repeats the model check every day. See Test your endpoint for the usual failures and fixes.
Capacity and heartbeats
Capacity tells BatchRouter how much work a lane can take. You can set it in the portal, or from your own systems through the API:
| Endpoint | What it does |
|---|---|
GET https://api.batchrouter.com/v1/provider-portal/offerings | Lists your lanes and their IDs |
POST https://api.batchrouter.com/v1/provider-portal/offerings/{offering_id}/capacity | Sets the lane's standing capacity |
POST https://api.batchrouter.com/v1/provider-portal/offerings/{offering_id}/heartbeat | Updates capacity and keeps it open for a time-to-live |
Authenticate with a BatchRouter API key from the account that owns the provider, and name the provider in a header:
curl -X POST "https://api.batchrouter.com/v1/provider-portal/offerings/$OFFERING_ID/heartbeat" \
-H "Authorization: Bearer $BATCHROUTER_API_KEY" \
-H "x-batchrouter-provider-id: your-provider-slug" \
-H "Content-Type: application/json" \
-d '{"available_queue_items": 5000, "heartbeat_ttl_seconds": 300}'Both endpoints take the capacity fields from
Publish capacity, such as max_batch_items,
max_concurrent_batches, available_queue_items and available_queue_tokens. A heartbeat also
takes heartbeat_ttl_seconds, from 30 to 3600. Without it, the lane's configured time-to-live
applies.
Once you send heartbeats, keep sending them
Each heartbeat sets an expiry: the lane's capacity is open until the time-to-live runs out. If the
heartbeats stop, the lane stops taking work. To go back to standing capacity, call /capacity with
"capacity_expires_at": null and "last_heartbeat_at": null.
Running on vLLM
vLLM serves GET /v1/models and POST /v1/chat/completions, so it covers direct requests and the
model check as it is. For batch work you run a batch front end in front of it that hosts the batch
endpoints on this page. vllm run-batch reads the same JSONL input format and writes the same output
line format, which makes it a possible starting point for that front end.