BatchRouter Docs

Evals, sweeps & quality-gated routing

Sweep one dataset across many provider lanes at once, get a scorecard with actual cost and a cost/quality Pareto frontier, then route with cheapest_passing so only lanes that cleared your bar are eligible.

BatchRouter routes to the cheapest eligible lane. The eval layer lets you turn that into a stronger claim: the cheapest lane that passes your quality bar, with the evidence to back it.

Because batch execution is cheap and latency-insensitive, the natural unit here is not one eval run but a sweep — the same dataset against many candidate lanes at once, priced and settled as a single batch.

Grading is deterministic and runs inside BatchRouter, so scoring costs you nothing. You pay only for the inference the sweep actually ran.

The loop

  1. Create a dataset — your inputs, plus expected outputs where you have them.
  2. Create an eval — a task prompt and one or more graders.
  3. Run a sweep — that dataset against N lanes, as one batch with one quote and one settlement.
  4. Read the scorecard — per-lane quality, actual cost, latency, errors, retries, and the cost/quality Pareto frontier.
  5. Route with cheapest_passing — only lanes that cleared your bar are eligible.
  6. Turn on shadow sampling — keep the scorecard fresh from live traffic, and catch regressions automatically.

1. Create a dataset

A dataset is inputs plus optional expected outputs. Inputs use the same shape as batch items, so anything you already send to POST /v1/batches works unchanged.

curl https://api.batchrouter.com/v1/datasets \
  -H "Authorization: Bearer $BATCHROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "invoice extraction v1",
    "items": [
      {
        "customer_item_id": "inv-1",
        "input": { "input": "Extract vendor and total from: ACME LTD, $412.50" },
        "expected_output": { "vendor": "ACME LTD", "total": 412.5 }
      }
    ]
  }'

Importing an existing suite

Already have evals elsewhere? Migration is a config change, not a rewrite. Pass raw rows under import and name the format:

{
  "name": "imported from openai evals",
  "source_format": "openai_evals",
  "import": {
    "rows": [
      {
        "item": {
          "id": "s-1",
          "input": "Extract the vendor.",
          "ground_truth": { "vendor": "ACME" }
        }
      }
    ]
  }
}

Both the modern { "item": { … } } data-source rows and classic { "input": […], "ideal": "…" } JSONL are accepted.

If a row can't be imported, the 400 names the offending row index rather than failing the whole upload with a generic message.

Privacy

A dataset carries a privacy_tier (standard · confidential · restricted), the same tier the rest of the platform uses. Routing hard-gates on it, so a restricted dataset only ever sweeps across restricted-capable lanes.

2. Create an eval

An eval is a task prompt plus graders. v1 covers structured extraction — schema conformance and field accuracy — with objective graders only. No LLM-as-judge.

curl https://api.batchrouter.com/v1/evals \
  -H "Authorization: Bearer $BATCHROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "invoice fields",
    "output_schema": {
      "type": "object",
      "required": ["vendor", "total"],
      "properties": { "vendor": { "type": "string" }, "total": { "type": "number" } }
    },
    "graders": [
      { "kind": "json_schema", "weight": 2 },
      { "kind": "normalized_match", "field_path": "vendor" },
      { "kind": "numeric_tolerance", "field_path": "total", "absolute_tolerance": 0.01 }
    ]
  }'
GraderWhat it checksScore
json_schemaThe whole output conforms to output_schema.pass/fail
exact_matchOne field equals the expected value exactly (objects and arrays compare structurally).pass/fail
normalized_matchSame, ignoring case, punctuation and whitespace by default.pass/fail
numeric_toleranceOne numeric field is within an absolute and/or relative tolerance.pass/fail
wer / cerWord- or character-level transcription error rate.continuous (1 - error rate)
embedding_similaritySemantic closeness of a free-form text field.continuous (the similarity)

Weights are relative — they don't need to sum to 1.

Schemas are checked in full, or refused

POST /v1/evals rejects an output_schema containing a JSON Schema keyword the grader cannot enforce (400 unsupported_output_schema, naming them). A passing json_schema grade therefore means every constraint in your schema was actually checked — never that some were quietly skipped. Supported: type, const, enum, required, properties, items, minItems/maxItems, minimum/maximum/exclusiveMinimum/exclusiveMaximum, multipleOf, minLength/maxLength, pattern, uniqueItems, additionalProperties, minProperties/maxProperties, and allOf/anyOf/oneOf/not.

A grader that can't run — a field matcher on a row with no expected output — is skipped and left out of the average, not scored zero. A lane that produced no output at all does score zero: a lane that fails is worse than one that answers imperfectly.

Beyond structured extraction

Multimodal input already works: image, document, audio and video attachments route like any other batch item, so "extract these fields from this invoice image across 12 lanes" needs nothing special. What varies is how the output is graded.

  • Transcription — grade with wer or cer rather than normalized_match, which scores a one-word slip identically to nonsense. max_error_rate decides pass/fail while the score stays continuous, so a 95%-correct transcript reads as 0.95.
  • Free-form text — embedding_similarity compares meaning instead of characters. It is the only grader that costs anything; identical strings are embedded once per sweep and batched, so it stays far below the cost of the inference being graded.
  • Generated media (images, audio, video as output) is not gradeable yet.
{
  "name": "call transcription",
  "graders": [
    { "kind": "wer", "field_path": "transcript", "max_error_rate": 0.15 },
    {
      "kind": "embedding_similarity",
      "field_path": "summary",
      "min_similarity": 0.85
    }
  ]
}

3. Run a sweep

One dataset × one eval × N lanes:

curl https://api.batchrouter.com/v1/sweeps \
  -H "Authorization: Bearer $BATCHROUTER_API_KEY" \
  -H "Idempotency-Key: $(uuidgen)" \
  -H "Content-Type: application/json" \
  -d '{
    "dataset_id": "ds_...",
    "eval_id": "ev_...",
    "lanes": ["openai:gpt-5.4-mini", "auto:cheapest_5"]
  }'

lanes accepts explicit lane identifiers (openai, openai:gpt-5.4-mini) and selectors (auto:cheapest_n, up to 12). A selector never re-selects a lane you named, so the example above means that lane plus the five cheapest others.

A sweep is one batch. Its items are the dataset × lane cross-product, so it comes back with one batch_id, one quote_id, one credit reservation, one settlement against actual token usage, one artifact set, one webhook and the same 24-hour SLA as any other batch. Poll batch_id exactly as you would normally.

dataset_items × lanes is capped at the 5 000 batch-item limit. Over it, the request is rejected with sweep_too_large before any credit is reserved, and the error carries the arithmetic.

Screen first, then promote

A sweep costs dataset_items × lanes, so the cheapest way to compare a lot of lanes is not to run the whole dataset against all of them. Screen with sample_size, then promote the survivors:

# 1. Screen: 25 rows across the 8 cheapest lanes = 200 items, not 4 000.
curl https://api.batchrouter.com/v1/sweeps ... -d '{
  "dataset_id": "ds_...", "eval_id": "ev_...",
  "lanes": ["auto:cheapest_8"], "sample_size": 25
}'

# 2. Promote: full dataset, only the lanes that cleared 0.8.
curl https://api.batchrouter.com/v1/sweeps ... -d '{
  "dataset_id": "ds_...", "eval_id": "ev_...",
  "promote_from": { "sweep_id": "sw_...", "min_score": 0.8 }
}'

Sampling is evenly spaced and deterministic — not a prefix (which would systematically miss whatever your dataset's tail contains) and not random (so a retried create prices identically). Promotion re-resolves lanes against the live catalogue, so an offering withdrawn since the screening sweep is caught rather than silently reused. If nothing cleared the bar you get 409 no_lanes_to_promote with each lane's score.

These stay two sweeps, deliberately. Staging inside a single sweep would need either a second batch — and therefore a second settlement — or dispatching items after creation, which batch items don't allow. Two sweeps keeps the guarantee that each one is a single quote and a single settlement.

Don't pay twice for the same question

Sweeps repeat themselves. You re-run a dataset after tweaking one lane, or promote a screening sweep to the full dataset, and most of the questions are ones BatchRouter has already answered.

When an item's request is byte-identical and the lane resolves to the same provider offering at the same version, the result is reused instead of dispatched. Reused items consume no provider tokens, so they cost nothing — the sweep settles only for what it actually ran.

It is on by default for sweeps. Turn it off per sweep:

curl https://api.batchrouter.com/v1/sweeps ... -d '{
  "dataset_id": "ds_...", "eval_id": "ev_...",
  "lanes": ["auto:cheapest_8"],
  "reuse_results": false
}'

What busts the cache — any of these means a fresh run:

  • the prompt, the model, or any sampling parameter changes (the fingerprint covers the whole request)
  • the provider ships a new version of the offering
  • a different organization asks the same question — the cache is scoped to your organization and never shared
  • the item is on the restricted privacy tier, which never participates

Reuse can't tell that a request was meant to be stochastic. Two identical high-temperature requests fingerprint the same even though their answers genuinely differ. If re-asking the same question is the point of your eval, set reuse_results: false.

The scorecard reports reused_item_count per lane, so a lane that came back cheap because it was mostly reused is legible rather than mysteriously good.

Give a screening sweep its own deadline

Sweeps default to the platform's 24-hour SLA. That's the wrong shape for screening — "which lanes are worth a full run" is not an answer you want tomorrow.

curl https://api.batchrouter.com/v1/sweeps ... -d '{
  "dataset_id": "ds_...", "eval_id": "ev_...",
  "lanes": ["auto:cheapest_8"], "sample_size": 25,
  "deadline_seconds": 3600
}'

A lane that can't deliver inside the window expires and is reported as such on the scorecard — which is itself a screening result, and usually a decisive one. The sweep still settles once, against whatever its items actually consumed.

deadline_seconds can be anything from 5 minutes up to the 24-hour SLA. You can ask for an answer sooner than the SLA, never later. Omit it and nothing changes.

4. Read the scorecard

curl https://api.batchrouter.com/v1/sweeps/sw_.../scorecard \
  -H "Authorization: Bearer $BATCHROUTER_API_KEY"
{
  "final": true,
  "lanes": [
    {
      "lane_key": "together:meta-llama/Llama-3.3-70B-Instruct-Turbo",
      "score": 0.82,
      "pass_rate": 0.74,
      "item_count": 25,
      "reused_item_count": 9,
      "actual_cost": { "currency": "usd", "amount": "0.0210" },
      "control_plane_fee": { "currency": "usd", "amount": "0.0400" },
      "billed_cost": { "currency": "usd", "amount": "0.0610" },
      "mean_latency_seconds": 41.2,
      "error_count": 0,
      "retry_count": 1,
      "on_pareto_frontier": true
    }
  ],
  "pareto_frontier": ["together:meta-llama/Llama-3.3-70B-Instruct-Turbo", "openai:gpt-5.4-mini"]
}
  • actual_cost is what the lane consumed, not what it was quoted.
  • reused_item_count is how many of the lane's items were served from an earlier sweep rather than dispatched. Those cost nothing, so read it next to actual_cost — a lane can look cheap because it was fast or because most of it was reused.
  • billed_cost is what you actually pay — actual_cost plus the control-plane fee. On a cheap-lane screening sweep the fee can be most of the bill, so comparing lanes on actual_cost alone understates every lane, and understates the cheapest ones most.
  • control_plane_fee is charged per provider execution, not per lane. A lane that needed more than one execution — a retry, a provider failover — carries more than one fee, which is why the lane above shows 0.0400 rather than a single $0.02. It is not a flat offset you can mentally add to every lane. (Note this is not the same count as retry_count, which sums per-item attempts; one work-unit execution covers many items.)
  • The Pareto frontier keeps only lanes nothing beats on both cost and quality — the set you actually have to choose between. It is ranked on billed_cost, so it reflects what you pay rather than inference alone; because the fee varies with executions, a lane that retried can legitimately fall off a frontier it would have made on inference cost. It's ordered cheapest first. Lanes that couldn't be scored are excluded rather than treated as scoring zero.
  • final is false while the sweep is still running; the totals and the frontier are a live preview.
  • Individual lanes land as they finish. You don't wait for the slowest lane: a lane is scored and becomes gateable as soon as all of its items are done. The Pareto frontier spans every lane, so it is only settled once the whole sweep completes.

5. Route with cheapest_passing

Once you hold a scorecard, gate your production routing on it:

{
  "routing_mode": "cheapest_passing",
  "eval_id": "ev_...",
  "min_score": 0.95,
  "items": [ … ]
}

Route to the cheapest eligible lane that cleared your bar on your eval. Works on POST /v1/quotes/model and POST /v1/batches.

The quote comes back with an eval_gate explaining the decision — every candidate lane, whether it was admitted, its observed score, the bar, and when that score was measured:

{
  "eval_gate": {
    "eval_id": "ev_...",
    "min_score": 0.95,
    "admitted_providers": ["openai"],
    "blocked_providers": ["together"],
    "lanes": [
      {
        "provider": "together",
        "lane_key": "together:meta-llama/Llama-3.3-70B-Instruct-Turbo",
        "status": "below_min_score",
        "score": 0.82,
        "computed_at": "2026-09-01T00:00:00.000Z",
        "reason": "Scored 0.8200 on eval ev_..., below the 0.95 bar."
      }
    ]
  }
}

Three behaviours worth knowing up front:

  • A lane you've never swept is blocked, not admitted (status: "no_scorecard"). Routing to an unmeasured lane because it happens to be cheap would defeat the point of the mode.
  • If nothing clears the bar, the request fails with 409 no_lane_passes_eval_gate and the full rationale. It never quietly falls back to cheapest — believing you're quality-gated when you aren't would be worse than an error.
  • The gate compares a confidence bound, not the raw average. A mean of 0.95 over 20 rows and over 2 000 rows are not the same claim, so the gate uses the 95% Wilson lower bound (score_lower_bound) and requires at least min_sample_size scored items (default 20).

How many rows a bar needs

Because the gate uses a lower bound, an observed score sitting exactly on the bar does not pass, and a perfect lane needs roughly 73 rows to clear min_score: 0.95. A lane scored on too few items reports insufficient_evidence — a different problem from below_min_score, and the fix is to sweep more rows rather than change lanes. Lower min_sample_size if you want to gate on thinner evidence, knowing what that costs you.

6. Keep it fresh with shadow sampling

Scorecards go stale: models get re-quantised, offerings get re-versioned, quality drifts. Shadow sampling mirrors a slice of your production traffic onto a stronger lane and grades the delta.

{
  "shadow_sampling": { "eval_id": "ev_...", "rate": 0.05, "lane": "openai:gpt-5.4" },
  "items": [ … ]
}

5% of the batch also runs on openai:gpt-5.4; the stronger lane's output becomes the reference your production lane is graded against, and the result folds straight back into the scorecard cheapest_passing reads. A lane that drifts starts failing its own gate — no sweep needed.

The shadow items ride on the same batch, so they're quoted, reserved and settled with everything else, at what they actually consume. Sampling is deterministic and evenly spaced, so an idempotent retry prices the same batch.

Limits & scope

  • One eval type in v1: structured extraction. Objective graders only — no LLM-as-judge, so open-ended quality judgements are still out of reach.
  • Generated media cannot be graded; only outputs addressable as JSON text.
  • dataset_items × lanes ≤ 5 000; auto:cheapest_n ≤ 12 lanes; shadow rate ≤ 0.5.
  • The scorecard is the product — BatchRouter deliberately doesn't ship tracing, prompt management, a playground or an eval dashboard.

On this page