Onboarding checklist
The seven steps a provider completes in the BatchRouter provider portal before going live, including exactly what the endpoint test sends.
After you register, the provider portal shows a setup checklist. BatchRouter computes the checklist from the same checks that Go live runs, so a step marked complete in the portal is complete for routing too. Each open step says what's missing and links to the fix.
The endpoints your server hosts, and exactly what BatchRouter sends to them, are in the Provider API.
Choose your integration
You pick the integration when you register. It decides what work your lanes receive:
| Integration | Batch work | Direct requests |
|---|---|---|
| OpenAI Batch API (recommended) | Yes | Yes, at the lane's direct price (step 2) |
| Chat completions only | No | Yes |
A chat-only lane gets no batch work. If you start chat-only, you can switch the lane to the OpenAI Batch API once your server hosts it.
1. Add your models
Add every model you serve. Set each model's Model ID to exactly the ID your server lists at
GET /v1/models (in vLLM, the --served-model-name). BatchRouter sends this ID, dots and all,
as model with every request.
When a model matches one in the BatchRouter catalog, also pick that catalog identity so customers who ask for the model can reach you. Picking a catalog identity doesn't change your Model ID. A model name that overlaps with a different catalog model is blocked. Pick the matching identity or use a unique name.
If you register an open-weight model, include its Hugging Face repo ID (for example
meta-llama/Llama-3.1-8B-Instruct) and its quantization (for example bf16, fp8 or int4).
2. Set your prices
Prices are in USD per 1 million tokens. BatchRouter adds its margin on top, so the price you set is what you earn.
- Every model needs a price above $0.
- Prices above $1,000 per 1 million tokens are rejected. This usually catches a price entered per 1K tokens or per token.
On an OpenAI Batch API lane, this is your batch price, and the lane also needs a direct price. The direct price is required: it's what you earn for direct requests on the lane, and each batch price must stay below it. On a chat-completions-only lane, the price you set is your direct price.
If your batch API returns results while a batch is still running, turn on partial results for the lane. Customers then see finished items before the whole batch ends. See Partial results.
3. Test your endpoint
BatchRouter calls your endpoint with your API key to confirm it can reach you. The test does no real work and costs nothing.
| Integration | What BatchRouter sends | What passes |
|---|---|---|
| Both | GET {base URL}/v1/models with Authorization: Bearer <your key> | An OpenAI-style model list, {"object": "list", "data": [{"id": "your-model-id"}]}, that lists the Model ID of every model you added, exactly as BatchRouter sends it in model |
| OpenAI Batch API | Also GET {base URL}/v1/batches?limit=1 with the same key | An OpenAI-style list, {"object": "list", "data": [...]}. An empty data array passes |
Base URL. Enter your https:// base URL without a trailing /v1. BatchRouter adds
/v1/... itself, and strips a trailing /v1 if you include one. Don't include query strings or
fragments.
API key. Create a dedicated, revocable key for BatchRouter and enter it in the portal, where it's stored encrypted. Never send API keys by email.
Changing your key or base URL resets this step, so run the test again afterwards. If it fails, the portal shows the URL it called, the HTTP status and a hint. The usual fixes:
401or403: your endpoint rejected the key. BatchRouter sends it asAuthorization: Bearer <key>, so update the key in your endpoint settings to match what your server accepts.429or503: counted as busy, not as a broken endpoint. Run the test again when you have capacity.- Model ID mismatch: the test names each ID it couldn't find. Make the Model ID match what your
server lists, or have your server list and accept both names (vLLM's
--served-model-nametakes several). - Not an OpenAI-style model list: return JSON with a
dataarray of objects that have anid. 404from/v1/batches: your server doesn't host the batch API yet. Host it, or register the lane as chat completions only.- Timeouts or connection errors: make sure the endpoint is reachable from the public internet over HTTPS.
4. Publish capacity
Tell BatchRouter how much work you can take. Adding a model creates a lane with safe starting limits, which you can tune:
| Setting | What it controls |
|---|---|
| Target / max batch items | The preferred and maximum number of items in one job BatchRouter sends you |
| Target / max batch tokens | The preferred and maximum estimated tokens per job. Leave blank if item count is enough |
| Max concurrent batches | How many jobs can run on the lane at the same time |
| Available queue items / tokens | Room BatchRouter may reserve right now. Items must be above 0 for the lane to take work. Leave tokens blank for no token cap |
Capacity can expire automatically. If you send capacity heartbeats, each one keeps the lane open for its time-to-live. If the heartbeats stop, the lane stops taking work until you publish capacity again. If you don't send heartbeats, turn auto-expiry off.
You can also set capacity and send heartbeats from your own systems through the API. See Capacity and heartbeats.
5. Declare data handling
Routing only sends work to providers that have declared how they handle customer data. Declare:
- Retention: how many days you keep prompts and outputs.
0means zero data retention. - Zero data retention: whether you support it.
- Training: whether you train on customer data.
- Countries: the country or countries you run inference in. You need at least one before you can go live. See Processing countries.
- Results: confirm that results go back to BatchRouter for delivery.
Customers can filter routing by privacy tier and data handling, so declare what you actually do.
Processing countries
You declare countries, not regions. Name every country you run inference in. Through the API the
field is processing_countries, and it takes lower-case codes such as de or us.
- The list has 44 countries, including every EU member, China (
cn) and Hong Kong (hk). gbis read asuk.globalandeuare not countries and are refused.
BatchRouter derives your region tags (supported_regions) from the countries: the countries
themselves, plus eu when all of them are EU members. Every lane you publish carries those tags.
You can't set them separately.
The countries decide which work can reach you:
- Work with no region restriction can always reach you.
- Region-restricted work reaches you only when the customer's
allowed_regionscovers every country you run in, by name or through a group such aseu. If you run in Germany and France, you geteuwork but not Germany-only work. - Customers can't name China or Hong Kong, so a provider that runs there only gets unrestricted work.
You can change your countries at any time. There is no review step.
The customer-facing guide explains the same matching from the other side: Processing countries and regions.
6. Add payout details
Add the bank account BatchRouter pays your earnings into. Payouts are made monthly in USD by bank transfer. See earnings and payouts for what to enter and how the cycle works.
7. Go live
When steps 1–6 are complete, Go live re-runs every check, including a fresh endpoint test (the model list, and for the OpenAI Batch API the batch list), and activates your provider. Routing to you starts automatically. There's no manual approval step. New providers start with conservative routing limits.
BatchRouter operators can switch a provider's routing off and on. If routing to your provider is switched off, the portal says so, and no work is sent to you until it is switched back on.
Once you're live, BatchRouter keeps re-testing your endpoint. If a later test fails, or a lane's capacity expires, the portal shows a warning and that lane stops receiving work until you fix it.
Need help?
Email hello@batchrouter.com with your provider name. Never include API keys or bank details in email.