BatchRouter Docs

How routing works

How BatchRouter scores provider lanes on price, SLA, capacity, privacy, and tools, selects one or more lanes per batch, and records rejection receipts.

BatchRouter is a control plane between your batch and many AI providers. You submit one job; BatchRouter decides which provider lane (or lanes) actually run it. That decision is computed at quote time, is explainable, and is recorded so you can audit why a lane was chosen or rejected.

This page explains the routing pipeline conceptually — what a "lane" is, how lanes are scored and selected, why one batch can split across several lanes, and what the rejection codes mean. For the request/response field details, see the interactive API reference.

Lanes: the unit of routing

A lane is one concrete way to run part of your batch: a specific provider, on a specific model offering version, with its own price, capacity, SLA window, region, privacy posture, and declared hosted tools.

The key idea is that one customer batch is not necessarily one lane. You keep a single batch_id and a single set of results, but internally BatchRouter may execute your items across several lanes. A quote therefore returns a set of quote_lanes: each lane is flagged selected or not, and every non-selected lane (filtered out, or eligible-looking but not chosen) carries a rejection_code, a customer-safe rejection_reason, and a rejection_receipt explaining why. After a batch settles, the dedicated rejected_lanes array appears on the billing and usage receipts.

You never address a lane directly. You express requirements (model, SLA, privacy, region, tools, price ceiling) and a routing_mode, and BatchRouter resolves those into lanes. This keeps your integration stable even as providers, prices, and capacity change underneath you.

The pipeline

Routing happens when you create a quote (POST /v1/quotes/model or POST /v1/quotes/workflow) and is re-checked when you submit the batch (POST /v1/batches). It runs in two stages: hard gates first, then scoring.

  1. Preflight validation. Before any provider is considered, BatchRouter validates the request shape: JSONL/item shape, file types, context-window fit, hosted-tool support, JSON-schema/runtime capability, and webhook configuration. Failures return 400 with the problem under error.details.preflight, so you fix them before spending credits.

  2. Hard gates (eligibility). Each candidate lane must pass every gate to be considered at all. A lane is gated out if the provider does not support the requested task type or model, lacks the rate-limit capacity to finish inside the window, fails the privacy/region/data-handling constraints, is paused or unhealthy, is priced above your max_price, or does not declare a required hosted tool. Lanes that fail a gate become non-selected lanes with a rejection code (see below).

  3. Scoring (ranking). Surviving lanes are ranked. For direct routes, price is the dominant factor after the gates — the default outcome is the cheapest eligible lane. Scoring also weighs declared and observed capacity, estimated completion time, historical SLA adherence, recent error/retry/timeout rates, result-validation success, and provider trust tier. For workflow products, workflow fit and validation history can outrank raw price.

  4. Selection. Your routing_mode shapes which ranked lane(s) win (see Routing modes). The selected lanes, the estimated price, the expected completion window, and a routing explanation come back in the quote.

  5. Splitting (when needed). If no single lane can finish the whole batch in time — or splitting lowers price or reduces risk and you allow it — BatchRouter assigns different items to different lanes under the same batch_id. See Multi-lane splitting.

Creating a quote is free. Credits are only reserved when you accept a quote by creating a batch, and the final charge is settled from actual provider token usage. Quote first, inspect the lanes and the explanation, then submit. See Quotes and pricing.

Routing modes

routing_mode (a field on the batch/quote request) tells the selector how to trade price against SLA, supply source, and privacy. The default is cheapest.

ModeWhat it optimizes for
cheapestLowest eligible price (the default).
sla_awareBalances price against SLA adherence and completion confidence.
public_onlyRestricts supply to public, first-party provider lanes.
edge_onlyRestricts supply to managed edge / BatchProviderApi lanes.
hybridMixes public and edge supply as scoring dictates.
privacy_constrainedPrioritizes lanes that satisfy stricter data-handling requirements.

Mode interacts with the other knobs rather than replacing them: cheapest still has to pass every hard gate, and privacy_constrained still ranks the privacy-eligible lanes by price and reliability.

The SLA and deadlines

Every batch has one SLA: a 24-hour completion window. That window is what the router has to honor, which in turn affects which lanes are eligible (a provider whose processing window can't finish inside the deadline is gated out).

sla_tier accepts standard, flex, and priority for compatibility. All three are treated as standard, so the value doesn't change the window or the price.

Batch work goes to lanes that take whole batches. A provider lane that serves chat completions only gets direct requests, not batch work.

A batch carries a 24-hour SLA deadline. Standard built-in lanes (e.g. OpenAI and Anthropic native batch) run on a 24-hour window; most other providers run on a tighter provider-side processing window staged inside your customer SLA, which leaves room for retries and fallback before your deadline. You can read the deadline back from GET /v1/batches/{batchId}.

Privacy tiers

privacy_tier constrains which providers are even eligible, based on the privacy tiers each provider declares:

  • standard — the default; any eligible provider.
  • confidential — routes only to providers that declare the confidential tier. The tier is the provider's own declaration, not a zero-data-retention guarantee. The router adds a zero-data-retention check for this tier only when routing_mode is privacy_constrained. In every other mode, a confidential batch can reach a provider that declares no zero data retention, and no built-in provider declares it today.
  • restricted — only providers that declare the restricted tier are eligible. Like every tier, it is the provider's own declaration, and no built-in provider declares it.

A lane whose provider does not declare the requested tier is rejected with privacy_tier_mismatch rather than silently downgraded. BatchRouter checks the provider's declaration; it cannot control what a provider does with a payload after dispatch.

Multi-lane splitting

Because your batch is decoupled from any single provider, BatchRouter can split execution across lanes while you keep one batch_id and one result set. Splitting happens when:

  • the cheapest lane can't finish the whole batch inside the SLA window,
  • you've allowed splitting and it lowers the total price,
  • splitting reduces execution risk for a large job, or
  • an accepted lane becomes unavailable before dispatch.

Mixed-model submissions split naturally: items targeting different models become model-specific lanes under the one batch. When you poll GET /v1/batches/{batchId}, per-lane statuses are exposed so you can see how each internal execution is progressing.

Rejection receipts

Every lane that was considered but not selected is recorded with a normalized, machine-readable rejection_code, a customer-safe rejection_reason, and a persisted rejection_receipt. This is what makes routing explainable: you can see why a cheaper or closer lane didn't win.

Common rejection codes include:

CodeMeaning
capacity_fullThe lane had no available capacity to reserve for your work.
context_window_exceededYour items don't fit the lane's model context window.
privacy_tier_mismatchThe lane can't satisfy the requested privacy_tier.
region_unavailableThe lane's provider doesn't match your allowed_regions. A provider that declares the countries it runs in must have every one of them covered by your list. A provider marked global always passes this check, so allowed_regions is not a residency guarantee. See Region matching.
missing_web_searchThe lane's provider offering doesn't declare a required hosted tool (here, web search); the route explanation surfaces this as a tool_support gate.
stale_heartbeatThe lane's live-capacity heartbeat was too stale to treat as routable supply.

If hard gates eliminate every lane — for example, no eligible lane fits under your max_price, or no lane declares a required tool — quote creation fails rather than routing to a lane that doesn't meet your requirements. The rejection receipts on the failed quote tell you which constraint to relax. See the full code list in the API reference.

Region matching

allowed_regions limits which providers can take your work. It accepts global, eu, and a fixed list of countries. Leave it out, or include global, and any provider can take the work.

When you restrict regions, the check depends on the provider:

  • Providers that declare countries. A registered provider declares the country or countries it runs inference in (processing_countries). It gets region-restricted work only when your list covers every one of those countries, by name or through a group such as eu. A provider in Germany and France takes ["eu"] work but not ["de"] work. A provider in Germany and China takes only unrestricted work.
  • Providers without declared countries. Built-in providers, and registered providers that have not declared countries yet, match when their region tags share a value with your list. A provider marked global always matches.

Built-in providers marked global still take region-restricted work, so allowed_regions is not an enforced residency guarantee.

Providers can run in countries that allowed_regions can't name, such as China (cn) and Hong Kong (hk). A provider that declares one of those countries only gets unrestricted work.

Hosted tools as a gate

If your items need a provider-hosted tool — web_search, python_execution, calculator, time, file_search, or retrieval — declare it with required_tools. BatchRouter merges quote-level and item-level requirements, infers supported tool declarations from your item payloads, and rejects any lane whose selected provider offering doesn't declare every required tool. The route explanation surfaces this as a tool_support gate with the required-vs-offered tool lists.

Fallback

Routing doesn't stop at the first dispatch. If a selected lane fails, BatchRouter applies the configured fallback policy to keep the batch inside its SLA: it can re-contract the same provider within the window, hand the work to the next-best eligible lane, or fall back to a designated fallback provider for the same quoted model set when policy allows. Fallback is bounded by your original requirements — a fallback lane still has to pass the same hard gates (price ceiling, privacy, region, tools), so it never quietly violates a constraint you set.

Next steps

On this page