> ## Documentation Index
> Fetch the complete documentation index at: https://docs.while.ai/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> Install with `uv add whileai`; import as `import whileai as wai`.
> Run the offline path first (`simulator=False`, `wai.seeded_agent`, a callable judge); no key is needed for it.
> Report every pass rate with its interval and n, as `scored.pass_at` prints it.

# A model router, measured against random routing

> Send each request to a cheap model or to the frontier model, and measure whether the router earns its keep.

<Note>The scripts are in the repository, not in the installed package. Clone it,
then `cd recipes/02-measure/model-router` before running the commands below. [Browse this recipe on GitHub](https://github.com/whilehq/whileai-sdk/tree/main/recipes/02-measure/model-router).</Note>

Send each request to a cheap model or to the frontier model, and measure
whether the router earns its keep. A router is only worth building if it
beats sending the same share of traffic to the frontier model at random, so
every router here is scored against that control, paired on held-out
questions, with a 95% interval.

What you will learn: why "the router keeps 90% of quality at 30% of the
cost" means nothing without the random control at the same share; that a
router reading only the prompt barely beats random on this traffic, while a
cascade that reads the cheap model's own answers does; and how large the
gap to a perfect router still is. You need an OpenAI-compatible key
(`OPENAI_API_KEY`; an OpenRouter key works with the default `--base-url`).
`--dry-run` needs nothing. About 12 minutes and \$5.71 live, 15 seconds offline.

## Run it

```bash theme={"theme":"vitesse-dark"}
uv add whileai openai
cd recipes/02-measure/model-router
python run.py --dry-run          # offline: seeded stand-in models, what smoke.sh runs
python run.py                    # live: 3,988 calls, about $6 on OpenRouter
python run.py --cheap openai/gpt-5-nano --strong anthropic/claude-opus-5.5 \
  --cheap-price 0.05,0.4 --strong-price 4,20   # your own pair
```

| flag | default | what it does |
| - | - | - |
| `--cheap` | `google/gemma-3-12b-it` | the model most traffic should land on |
| `--strong` | `anthropic/claude-sonnet-5.5` | the frontier model it replaces |
| `--cheap-price`, `--strong-price` | `0.05,0.15`, `2,10` | USD per million tokens, in and out, for the cost column |
| `--base-url` | `$OPENAI_BASE_URL`, else OpenRouter | any OpenAI-compatible server |
| `--concurrency` | 16 | calls in flight |
| `--margin` | 0.02 | how far below all-frontier still counts as "kept" |
| `--limit` | all 1,190 | fewer questions, for a smoke run |
| `--fresh` | off | ignore `out/samples.jsonl` and call again |
| `--dry-run` | off | no key, no calls |

## The setup

**Traffic.** 1,190 MMLU-Pro test questions (Wang et al. 2024 \[1]), 85 from
each of 14 subjects, frozen in `tasks.jsonl` by `make_tasks.py`. They stand
in for an internal assistant's mixed traffic: law, business, health,
engineering and so on, ten options each, graded by `wai.verify.MultipleChoice`
with no judge. A third of each subject is held out by a hash of the question
id (772 train, 418 held out). Nothing a router learns comes from the held-out
third.

**Models.** Gemma 3 12B (open weights, one GPU) and Claude Sonnet 5.5, both
through OpenRouter at temperature 0.7. The cheap model answers every question
twice. The frontier model answers train once and held-out twice, so every
number below is read on a second, independent draw too.

**Routers.** Each sends the top share of held-out questions to the frontier
model:

* `subject`: the free rule. Escalate the subjects where the frontier model
  gained most on train.
* `text`: logistic regression on hashed word n-grams of the question,
  predicting "the cheap model misses" from its two train draws (the win-label
  router of RouteLLM \[2], in numpy).
* `cascade`: answer with the cheap model twice; escalate when the two
  answers disagree (FrugalGPT \[3], with self-consistency \[4] as the
  confidence). It pays for two cheap calls on every question, and its share
  is whatever the disagreement rate turns out to be.
* `oracle`: escalate exactly the questions where the frontier model is right
  and the cheap one is wrong. The ceiling, not a router.

**Controls.** `vs random` compares each router with random routing at the
same frontier share, written as 40 rollouts per question so the control is
exact rather than one lucky draw. `vs all frontier` compares it with the
baseline it is trying to replace. Both come from `wai.compare`, paired by
question.

## Result

Run 2026-10-01, every call made in that run, nothing reused. Held out, draw 1:

| | frontier share | accuracy | cost vs all-frontier | vs random at same share | vs all frontier |
| - | - | - | - | - | - |
| all frontier | 1.00 | 0.766 | 1.00x | | |
| all cheap | 0.00 | 0.505 | 0.01x | | −0.261 \[−0.316, −0.208] |
| `text` | 0.40 | 0.632 | 0.46x | +0.022 \[−0.007, +0.054] | −0.134 \[−0.177, −0.091] |
| `subject` | 0.40 | 0.656 | 0.50x | +0.046 \[+0.015, +0.076] | −0.110 \[−0.153, −0.067] |
| `text` | 0.26 | 0.586 | 0.35x | +0.016 \[−0.010, +0.043] | −0.179 \[−0.227, −0.132] |
| **`cascade`** | **0.26** | **0.624** | **0.40x** | **+0.054 \[+0.024, +0.084]** | −0.141 \[−0.187, −0.098] |
| `oracle` | 0.33 | 0.835 | 0.39x | +0.245 \[+0.215, +0.275] | **+0.069 \[+0.048, +0.093]** |

`python run.py` prints all shares (0.2, 0.4, 0.6) and both draws;
`results.json` has every number.

**The cascade is the only router that beats random on both draws:** +5.4
points \[+2.4, +8.4] on draw 1, +4.5 \[+1.4, +7.5] on draw 2. It is the only
router that reads the cheap model's answers rather than the question.

**The prompt-only routers are about the same as random.** The text router's
interval crosses zero at every share on both draws, and its AUC for "the cheap
model will miss" is 0.588. The subject rule clears random on draw 1 (+4.6 at
40%) and does not replicate on draw 2 (+2.5 \[−0.1, +5.4]). What a question
looks like says little about whether a 12B model will get it wrong.

**The room is real.** The oracle sends 33% of traffic to the frontier model
and beats all-frontier by 6.9 points \[+4.8, +9.3] at 0.39x the cost, because
on about 7% of questions the cheap model is right and the frontier model is
wrong. A router that found the third of questions that need the frontier model
would win on both quality and cost. None of the three here comes close.

**No router keeps quality within 2 points of all-frontier** below 95% frontier
share. Read that as the honest answer to "how much traffic can we move off the
frontier model for free" on this traffic: almost none, with these routers.

### Limits

* **MMLU-Pro is a stand-in for your traffic, and a public one.** Either model
  may have seen the questions. Real enterprise traffic has more easy
  questions, which helps every router.
* **Costs are list prices times tokens used.** Escalated questions cost more
  than average (they are longer and get longer answers), so a 0.26 share costs
  0.40x, not 0.26x.
* **One train split, one router fit.** The second draw replicates the
  evaluation, not the router's training.
* **The cascade has one operating point.** Disagreement between two draws
  gives one share. More cheap draws or an agreement threshold would trace a
  curve; that is the next run.

## Next

Swap `tasks.jsonl` for a sample of your own traffic and the verifier for one
that grades it; the routers and the controls do not change. Before you build
a router, run `--dry-run` with your pair, then live with `--limit 200`: if
`cascade` does not beat random on 200 of your questions, a router will not
pay for itself.

## References

1. Wang et al. 2024. *MMLU-Pro: A More Robust and Challenging Multi-Task
   Language Understanding Benchmark.* arXiv:2406.01574.
2. Ong et al. 2024. *RouteLLM: Learning to Route LLMs with Preference Data.*
   arXiv:2406.18665.
3. Chen, Zaharia, Zou 2023. *FrugalGPT: How to Use Large Language Models
   While Reducing Cost and Improving Performance.* arXiv:2305.05176.
4. Wang et al. 2022. *Self-Consistency Improves Chain of Thought Reasoning in
   Language Models.* arXiv:2203.11171.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.