BatchRouter Docs

Onboarding checklist

The seven steps a provider completes in the BatchRouter provider portal before going live, including exactly what the endpoint test sends.

After you register, the provider portal shows a setup checklist. BatchRouter computes the checklist from the same checks that Go live runs, so a step marked complete in the portal is complete for routing too. Each open step says what's missing and links to the fix.

The endpoints your server hosts, and exactly what BatchRouter sends to them, are in the Provider API.

Choose your integration

You pick the integration when you register. It decides what work your lanes receive:

IntegrationBatch workDirect requests
OpenAI Batch API (recommended)YesYes, at the lane's direct price (step 2)
Chat completions onlyNoYes

A chat-only lane gets no batch work. If you start chat-only, you can switch the lane to the OpenAI Batch API once your server hosts it.

1. Add your models

Add every model you serve. Set each model's Model ID to exactly the ID your server lists at GET /v1/models (in vLLM, the --served-model-name). BatchRouter sends this ID, dots and all, as model with every request.

When a model matches one in the BatchRouter catalog, also pick that catalog identity so customers who ask for the model can reach you. Picking a catalog identity doesn't change your Model ID. A model name that overlaps with a different catalog model is blocked. Pick the matching identity or use a unique name.

If you register an open-weight model, include its Hugging Face repo ID (for example meta-llama/Llama-3.1-8B-Instruct) and its quantization (for example bf16, fp8 or int4).

2. Set your prices

Prices are in USD per 1 million tokens. BatchRouter adds its margin on top, so the price you set is what you earn.

  • Every model needs a price above $0.
  • Prices above $1,000 per 1 million tokens are rejected. This usually catches a price entered per 1K tokens or per token.

On an OpenAI Batch API lane, this is your batch price, and the lane also needs a direct price. The direct price is required: it's what you earn for direct requests on the lane, and each batch price must stay below it. On a chat-completions-only lane, the price you set is your direct price.

If your batch API returns results while a batch is still running, turn on partial results for the lane. Customers then see finished items before the whole batch ends. See Partial results.

3. Test your endpoint

BatchRouter calls your endpoint with your API key to confirm it can reach you. The test does no real work and costs nothing.

IntegrationWhat BatchRouter sendsWhat passes
BothGET {base URL}/v1/models with Authorization: Bearer <your key>An OpenAI-style model list, {"object": "list", "data": [{"id": "your-model-id"}]}, that lists the Model ID of every model you added, exactly as BatchRouter sends it in model
OpenAI Batch APIAlso GET {base URL}/v1/batches?limit=1 with the same keyAn OpenAI-style list, {"object": "list", "data": [...]}. An empty data array passes

Base URL. Enter your https:// base URL without a trailing /v1. BatchRouter adds /v1/... itself, and strips a trailing /v1 if you include one. Don't include query strings or fragments.

API key. Create a dedicated, revocable key for BatchRouter and enter it in the portal, where it's stored encrypted. Never send API keys by email.

Changing your key or base URL resets this step, so run the test again afterwards. If it fails, the portal shows the URL it called, the HTTP status and a hint. The usual fixes:

  • 401 or 403: your endpoint rejected the key. BatchRouter sends it as Authorization: Bearer <key>, so update the key in your endpoint settings to match what your server accepts.
  • 429 or 503: counted as busy, not as a broken endpoint. Run the test again when you have capacity.
  • Model ID mismatch: the test names each ID it couldn't find. Make the Model ID match what your server lists, or have your server list and accept both names (vLLM's --served-model-name takes several).
  • Not an OpenAI-style model list: return JSON with a data array of objects that have an id.
  • 404 from /v1/batches: your server doesn't host the batch API yet. Host it, or register the lane as chat completions only.
  • Timeouts or connection errors: make sure the endpoint is reachable from the public internet over HTTPS.

4. Publish capacity

Tell BatchRouter how much work you can take. Adding a model creates a lane with safe starting limits, which you can tune:

SettingWhat it controls
Target / max batch itemsThe preferred and maximum number of items in one job BatchRouter sends you
Target / max batch tokensThe preferred and maximum estimated tokens per job. Leave blank if item count is enough
Max concurrent batchesHow many jobs can run on the lane at the same time
Available queue items / tokensRoom BatchRouter may reserve right now. Items must be above 0 for the lane to take work. Leave tokens blank for no token cap

Capacity can expire automatically. If you send capacity heartbeats, each one keeps the lane open for its time-to-live. If the heartbeats stop, the lane stops taking work until you publish capacity again. If you don't send heartbeats, turn auto-expiry off.

You can also set capacity and send heartbeats from your own systems through the API. See Capacity and heartbeats.

5. Declare data handling

Routing only sends work to providers that have declared how they handle customer data. Declare:

  • Retention: how many days you keep prompts and outputs. 0 means zero data retention.
  • Zero data retention: whether you support it.
  • Training: whether you train on customer data.
  • Countries: the country or countries you run inference in. You need at least one before you can go live. See Processing countries.
  • Results: confirm that results go back to BatchRouter for delivery.

Customers can filter routing by privacy tier and data handling, so declare what you actually do.

Processing countries

You declare countries, not regions. Name every country you run inference in. Through the API the field is processing_countries, and it takes lower-case codes such as de or us.

  • The list has 44 countries, including every EU member, China (cn) and Hong Kong (hk).
  • gb is read as uk.
  • global and eu are not countries and are refused.

BatchRouter derives your region tags (supported_regions) from the countries: the countries themselves, plus eu when all of them are EU members. Every lane you publish carries those tags. You can't set them separately.

The countries decide which work can reach you:

  • Work with no region restriction can always reach you.
  • Region-restricted work reaches you only when the customer's allowed_regions covers every country you run in, by name or through a group such as eu. If you run in Germany and France, you get eu work but not Germany-only work.
  • Customers can't name China or Hong Kong, so a provider that runs there only gets unrestricted work.

You can change your countries at any time. There is no review step.

The customer-facing guide explains the same matching from the other side: Processing countries and regions.

6. Add payout details

Add the bank account BatchRouter pays your earnings into. Payouts are made monthly in USD by bank transfer. See earnings and payouts for what to enter and how the cycle works.

7. Go live

When steps 1–6 are complete, Go live re-runs every check, including a fresh endpoint test (the model list, and for the OpenAI Batch API the batch list), and activates your provider. Routing to you starts automatically. There's no manual approval step. New providers start with conservative routing limits.

BatchRouter operators can switch a provider's routing off and on. If routing to your provider is switched off, the portal says so, and no work is sent to you until it is switched back on.

Once you're live, BatchRouter keeps re-testing your endpoint. If a later test fails, or a lane's capacity expires, the portal shows a warning and that lane stops receiving work until you fix it.

Need help?

Email hello@batchrouter.com with your provider name. Never include API keys or bank details in email.

On this page