Evals, sweeps & quality-gated routing
Sweep one dataset across many provider lanes at once, get a scorecard with actual cost and a cost/quality Pareto frontier, then route with cheapest_passing so only lanes that cleared your bar are eligible.
BatchRouter routes to the cheapest eligible lane. The eval layer lets you turn that into a stronger claim: the cheapest lane that passes your quality bar, with the evidence to back it.
Because batch execution is cheap and latency-insensitive, the natural unit here is not one eval run but a sweep — the same dataset against many candidate lanes at once, priced and settled as a single batch.
Grading is deterministic and runs inside BatchRouter, so scoring costs you nothing. You pay only for the inference the sweep actually ran.
The loop
- Create a dataset — your inputs, plus expected outputs where you have them.
- Create an eval — a task prompt and one or more graders.
- Run a sweep — that dataset against N lanes, as one batch with one quote and one settlement.
- Read the scorecard — per-lane quality, actual cost, latency, errors, retries, and the cost/quality Pareto frontier.
- Route with
cheapest_passing— only lanes that cleared your bar are eligible. - Turn on shadow sampling — keep the scorecard fresh from live traffic, and catch regressions automatically.
1. Create a dataset
A dataset is inputs plus optional expected outputs. Inputs use the same shape as batch items, so anything you already send to POST /v1/batches works unchanged.
curl https://api.batchrouter.com/v1/datasets \
-H "Authorization: Bearer $BATCHROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "invoice extraction v1",
"items": [
{
"customer_item_id": "inv-1",
"input": { "input": "Extract vendor and total from: ACME LTD, $412.50" },
"expected_output": { "vendor": "ACME LTD", "total": 412.5 }
}
]
}'Importing an existing suite
Already have evals elsewhere? Migration is a config change, not a rewrite. Pass raw rows under import and name the format:
{
"name": "imported from openai evals",
"source_format": "openai_evals",
"import": {
"rows": [
{
"item": {
"id": "s-1",
"input": "Extract the vendor.",
"ground_truth": { "vendor": "ACME" }
}
}
]
}
}Both the modern { "item": { … } } data-source rows and classic { "input": […], "ideal": "…" } JSONL are accepted.
If a row can't be imported, the 400 names the offending row index rather than failing the whole upload with a generic message.
Privacy
A dataset carries a privacy_tier (standard · confidential ·
restricted), the same tier the rest of the platform uses. Routing hard-gates
on it, so a restricted dataset only ever sweeps across restricted-capable
lanes.
2. Create an eval
An eval is a task prompt plus graders. v1 covers structured extraction — schema conformance and field accuracy — with objective graders only. No LLM-as-judge.
curl https://api.batchrouter.com/v1/evals \
-H "Authorization: Bearer $BATCHROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "invoice fields",
"output_schema": {
"type": "object",
"required": ["vendor", "total"],
"properties": { "vendor": { "type": "string" }, "total": { "type": "number" } }
},
"graders": [
{ "kind": "json_schema", "weight": 2 },
{ "kind": "normalized_match", "field_path": "vendor" },
{ "kind": "numeric_tolerance", "field_path": "total", "absolute_tolerance": 0.01 }
]
}'| Grader | What it checks | Score |
|---|---|---|
json_schema | The whole output conforms to output_schema. | pass/fail |
exact_match | One field equals the expected value exactly (objects and arrays compare structurally). | pass/fail |
normalized_match | Same, ignoring case, punctuation and whitespace by default. | pass/fail |
numeric_tolerance | One numeric field is within an absolute and/or relative tolerance. | pass/fail |
wer / cer | Word- or character-level transcription error rate. | continuous (1 - error rate) |
embedding_similarity | Semantic closeness of a free-form text field. | continuous (the similarity) |
Weights are relative — they don't need to sum to 1.
Schemas are checked in full, or refused
POST /v1/evals rejects an output_schema containing a JSON Schema
keyword the grader cannot enforce (400 unsupported_output_schema, naming
them). A passing json_schema grade therefore means every constraint in your
schema was actually checked — never that some were quietly skipped. Supported:
type, const, enum, required, properties, items,
minItems/maxItems,
minimum/maximum/exclusiveMinimum/exclusiveMaximum, multipleOf,
minLength/maxLength, pattern, uniqueItems, additionalProperties,
minProperties/maxProperties, and allOf/anyOf/oneOf/not.
A grader that can't run — a field matcher on a row with no expected output — is skipped and left out of the average, not scored zero. A lane that produced no output at all does score zero: a lane that fails is worse than one that answers imperfectly.
Beyond structured extraction
Multimodal input already works: image, document, audio and video attachments route like any other batch item, so "extract these fields from this invoice image across 12 lanes" needs nothing special. What varies is how the output is graded.
- Transcription — grade with
werorcerrather thannormalized_match, which scores a one-word slip identically to nonsense.max_error_ratedecides pass/fail while the score stays continuous, so a 95%-correct transcript reads as 0.95. - Free-form text —
embedding_similaritycompares meaning instead of characters. It is the only grader that costs anything; identical strings are embedded once per sweep and batched, so it stays far below the cost of the inference being graded. - Generated media (images, audio, video as output) is not gradeable yet.
{
"name": "call transcription",
"graders": [
{ "kind": "wer", "field_path": "transcript", "max_error_rate": 0.15 },
{
"kind": "embedding_similarity",
"field_path": "summary",
"min_similarity": 0.85
}
]
}3. Run a sweep
One dataset × one eval × N lanes:
curl https://api.batchrouter.com/v1/sweeps \
-H "Authorization: Bearer $BATCHROUTER_API_KEY" \
-H "Idempotency-Key: $(uuidgen)" \
-H "Content-Type: application/json" \
-d '{
"dataset_id": "ds_...",
"eval_id": "ev_...",
"lanes": ["openai:gpt-5.4-mini", "auto:cheapest_5"]
}'lanes accepts explicit lane identifiers (openai, openai:gpt-5.4-mini) and selectors (auto:cheapest_n, up to 12). A selector never re-selects a lane you named, so the example above means that lane plus the five cheapest others.
A sweep is one batch. Its items are the dataset × lane cross-product, so it comes back with one batch_id, one quote_id, one credit reservation, one settlement against actual token usage, one artifact set, one webhook and the same 24-hour SLA as any other batch. Poll batch_id exactly as you would normally.
dataset_items × lanes is capped at the 5 000 batch-item limit. Over it, the
request is rejected with sweep_too_large before any credit is reserved,
and the error carries the arithmetic.
Screen first, then promote
A sweep costs dataset_items × lanes, so the cheapest way to compare a lot of lanes is not to
run the whole dataset against all of them. Screen with sample_size, then promote the survivors:
# 1. Screen: 25 rows across the 8 cheapest lanes = 200 items, not 4 000.
curl https://api.batchrouter.com/v1/sweeps ... -d '{
"dataset_id": "ds_...", "eval_id": "ev_...",
"lanes": ["auto:cheapest_8"], "sample_size": 25
}'
# 2. Promote: full dataset, only the lanes that cleared 0.8.
curl https://api.batchrouter.com/v1/sweeps ... -d '{
"dataset_id": "ds_...", "eval_id": "ev_...",
"promote_from": { "sweep_id": "sw_...", "min_score": 0.8 }
}'Sampling is evenly spaced and deterministic — not a prefix (which would systematically miss
whatever your dataset's tail contains) and not random (so a retried create prices identically).
Promotion re-resolves lanes against the live catalogue, so an offering withdrawn since the
screening sweep is caught rather than silently reused. If nothing cleared the bar you get
409 no_lanes_to_promote with each lane's score.
These stay two sweeps, deliberately. Staging inside a single sweep would need either a second batch — and therefore a second settlement — or dispatching items after creation, which batch items don't allow. Two sweeps keeps the guarantee that each one is a single quote and a single settlement.
Don't pay twice for the same question
Sweeps repeat themselves. You re-run a dataset after tweaking one lane, or promote a screening sweep to the full dataset, and most of the questions are ones BatchRouter has already answered.
When an item's request is byte-identical and the lane resolves to the same provider offering at the same version, the result is reused instead of dispatched. Reused items consume no provider tokens, so they cost nothing — the sweep settles only for what it actually ran.
It is on by default for sweeps. Turn it off per sweep:
curl https://api.batchrouter.com/v1/sweeps ... -d '{
"dataset_id": "ds_...", "eval_id": "ev_...",
"lanes": ["auto:cheapest_8"],
"reuse_results": false
}'What busts the cache — any of these means a fresh run:
- the prompt, the model, or any sampling parameter changes (the fingerprint covers the whole request)
- the provider ships a new version of the offering
- a different organization asks the same question — the cache is scoped to your organization and never shared
- the item is on the
restrictedprivacy tier, which never participates
Reuse can't tell that a request was meant to be stochastic. Two identical
high-temperature requests fingerprint the same even though their answers
genuinely differ. If re-asking the same question is the point of your eval,
set reuse_results: false.
The scorecard reports reused_item_count per lane, so a lane that came back cheap because it was
mostly reused is legible rather than mysteriously good.
Give a screening sweep its own deadline
Sweeps default to the platform's 24-hour SLA. That's the wrong shape for screening — "which lanes are worth a full run" is not an answer you want tomorrow.
curl https://api.batchrouter.com/v1/sweeps ... -d '{
"dataset_id": "ds_...", "eval_id": "ev_...",
"lanes": ["auto:cheapest_8"], "sample_size": 25,
"deadline_seconds": 3600
}'A lane that can't deliver inside the window expires and is reported as such on the scorecard — which is itself a screening result, and usually a decisive one. The sweep still settles once, against whatever its items actually consumed.
deadline_seconds can be anything from 5 minutes up to the 24-hour SLA. You can ask for an answer
sooner than the SLA, never later. Omit it and nothing changes.
4. Read the scorecard
curl https://api.batchrouter.com/v1/sweeps/sw_.../scorecard \
-H "Authorization: Bearer $BATCHROUTER_API_KEY"{
"final": true,
"lanes": [
{
"lane_key": "together:meta-llama/Llama-3.3-70B-Instruct-Turbo",
"score": 0.82,
"pass_rate": 0.74,
"item_count": 25,
"reused_item_count": 9,
"actual_cost": { "currency": "usd", "amount": "0.0210" },
"control_plane_fee": { "currency": "usd", "amount": "0.0400" },
"billed_cost": { "currency": "usd", "amount": "0.0610" },
"mean_latency_seconds": 41.2,
"error_count": 0,
"retry_count": 1,
"on_pareto_frontier": true
}
],
"pareto_frontier": ["together:meta-llama/Llama-3.3-70B-Instruct-Turbo", "openai:gpt-5.4-mini"]
}actual_costis what the lane consumed, not what it was quoted.reused_item_countis how many of the lane's items were served from an earlier sweep rather than dispatched. Those cost nothing, so read it next toactual_cost— a lane can look cheap because it was fast or because most of it was reused.billed_costis what you actually pay —actual_costplus the control-plane fee. On a cheap-lane screening sweep the fee can be most of the bill, so comparing lanes onactual_costalone understates every lane, and understates the cheapest ones most.control_plane_feeis charged per provider execution, not per lane. A lane that needed more than one execution — a retry, a provider failover — carries more than one fee, which is why the lane above shows0.0400rather than a single$0.02. It is not a flat offset you can mentally add to every lane. (Note this is not the same count asretry_count, which sums per-item attempts; one work-unit execution covers many items.)- The Pareto frontier keeps only lanes nothing beats on both cost and quality — the set you actually have to choose between. It is ranked on
billed_cost, so it reflects what you pay rather than inference alone; because the fee varies with executions, a lane that retried can legitimately fall off a frontier it would have made on inference cost. It's ordered cheapest first. Lanes that couldn't be scored are excluded rather than treated as scoring zero. finalisfalsewhile the sweep is still running; the totals and the frontier are a live preview.- Individual lanes land as they finish. You don't wait for the slowest lane: a lane is scored and becomes gateable as soon as all of its items are done. The Pareto frontier spans every lane, so it is only settled once the whole sweep completes.
5. Route with cheapest_passing
Once you hold a scorecard, gate your production routing on it:
{
"routing_mode": "cheapest_passing",
"eval_id": "ev_...",
"min_score": 0.95,
"items": [ … ]
}Route to the cheapest eligible lane that cleared your bar on your eval. Works on POST /v1/quotes/model and POST /v1/batches.
The quote comes back with an eval_gate explaining the decision — every candidate lane, whether it was admitted, its observed score, the bar, and when that score was measured:
{
"eval_gate": {
"eval_id": "ev_...",
"min_score": 0.95,
"admitted_providers": ["openai"],
"blocked_providers": ["together"],
"lanes": [
{
"provider": "together",
"lane_key": "together:meta-llama/Llama-3.3-70B-Instruct-Turbo",
"status": "below_min_score",
"score": 0.82,
"computed_at": "2026-09-01T00:00:00.000Z",
"reason": "Scored 0.8200 on eval ev_..., below the 0.95 bar."
}
]
}
}Three behaviours worth knowing up front:
- A lane you've never swept is blocked, not admitted (
status: "no_scorecard"). Routing to an unmeasured lane because it happens to be cheap would defeat the point of the mode. - If nothing clears the bar, the request fails with
409 no_lane_passes_eval_gateand the full rationale. It never quietly falls back tocheapest— believing you're quality-gated when you aren't would be worse than an error. - The gate compares a confidence bound, not the raw average. A mean of 0.95 over 20 rows and over 2 000 rows are not the same claim, so the gate uses the 95% Wilson lower bound (
score_lower_bound) and requires at leastmin_sample_sizescored items (default 20).
How many rows a bar needs
Because the gate uses a lower bound, an observed score sitting exactly on
the bar does not pass, and a perfect lane needs roughly 73 rows to
clear min_score: 0.95. A lane scored on too few items reports
insufficient_evidence — a different problem from below_min_score, and the
fix is to sweep more rows rather than change lanes. Lower min_sample_size if
you want to gate on thinner evidence, knowing what that costs you.
6. Keep it fresh with shadow sampling
Scorecards go stale: models get re-quantised, offerings get re-versioned, quality drifts. Shadow sampling mirrors a slice of your production traffic onto a stronger lane and grades the delta.
{
"shadow_sampling": { "eval_id": "ev_...", "rate": 0.05, "lane": "openai:gpt-5.4" },
"items": [ … ]
}5% of the batch also runs on openai:gpt-5.4; the stronger lane's output becomes the reference your production lane is graded against, and the result folds straight back into the scorecard cheapest_passing reads. A lane that drifts starts failing its own gate — no sweep needed.
The shadow items ride on the same batch, so they're quoted, reserved and settled with everything else, at what they actually consume. Sampling is deterministic and evenly spaced, so an idempotent retry prices the same batch.
Limits & scope
- One eval type in v1: structured extraction. Objective graders only — no LLM-as-judge, so open-ended quality judgements are still out of reach.
- Generated media cannot be graded; only outputs addressable as JSON text.
dataset_items × lanes≤ 5 000;auto:cheapest_n≤ 12 lanes; shadowrate≤ 0.5.- The scorecard is the product — BatchRouter deliberately doesn't ship tracing, prompt management, a playground or an eval dashboard.