The scripts are in the repository, not in the installed package. Clone it,
then
cd recipes/02-measure/model-router before running the commands below. Browse this recipe on GitHub.OPENAI_API_KEY; an OpenRouter key works with the default --base-url).
--dry-run needs nothing. About 12 minutes and $5.71 live, 15 seconds offline.
Run it
The setup
Traffic. 1,190 MMLU-Pro test questions (Wang et al. 2024 [1]), 85 from each of 14 subjects, frozen intasks.jsonl by make_tasks.py. They stand
in for an internal assistant’s mixed traffic: law, business, health,
engineering and so on, ten options each, graded by wai.verify.MultipleChoice
with no judge. A third of each subject is held out by a hash of the question
id (772 train, 418 held out). Nothing a router learns comes from the held-out
third.
Models. Gemma 3 12B (open weights, one GPU) and Claude Sonnet 5.5, both
through OpenRouter at temperature 0.7. The cheap model answers every question
twice. The frontier model answers train once and held-out twice, so every
number below is read on a second, independent draw too.
Routers. Each sends the top share of held-out questions to the frontier
model:
subject: the free rule. Escalate the subjects where the frontier model gained most on train.text: logistic regression on hashed word n-grams of the question, predicting “the cheap model misses” from its two train draws (the win-label router of RouteLLM [2], in numpy).cascade: answer with the cheap model twice; escalate when the two answers disagree (FrugalGPT [3], with self-consistency [4] as the confidence). It pays for two cheap calls on every question, and its share is whatever the disagreement rate turns out to be.oracle: escalate exactly the questions where the frontier model is right and the cheap one is wrong. The ceiling, not a router.
vs random compares each router with random routing at the
same frontier share, written as 40 rollouts per question so the control is
exact rather than one lucky draw. vs all frontier compares it with the
baseline it is trying to replace. Both come from wai.compare, paired by
question.
Result
Run 2026-10-01, every call made in that run, nothing reused. Held out, draw 1:python run.py prints all shares (0.2, 0.4, 0.6) and both draws;
results.json has every number.
The cascade is the only router that beats random on both draws: +5.4
points [+2.4, +8.4] on draw 1, +4.5 [+1.4, +7.5] on draw 2. It is the only
router that reads the cheap model’s answers rather than the question.
The prompt-only routers are about the same as random. The text router’s
interval crosses zero at every share on both draws, and its AUC for “the cheap
model will miss” is 0.588. The subject rule clears random on draw 1 (+4.6 at
40%) and does not replicate on draw 2 (+2.5 [−0.1, +5.4]). What a question
looks like says little about whether a 12B model will get it wrong.
The room is real. The oracle sends 33% of traffic to the frontier model
and beats all-frontier by 6.9 points [+4.8, +9.3] at 0.39x the cost, because
on about 7% of questions the cheap model is right and the frontier model is
wrong. A router that found the third of questions that need the frontier model
would win on both quality and cost. None of the three here comes close.
No router keeps quality within 2 points of all-frontier below 95% frontier
share. Read that as the honest answer to “how much traffic can we move off the
frontier model for free” on this traffic: almost none, with these routers.
Limits
- MMLU-Pro is a stand-in for your traffic, and a public one. Either model may have seen the questions. Real enterprise traffic has more easy questions, which helps every router.
- Costs are list prices times tokens used. Escalated questions cost more than average (they are longer and get longer answers), so a 0.26 share costs 0.40x, not 0.26x.
- One train split, one router fit. The second draw replicates the evaluation, not the router’s training.
- The cascade has one operating point. Disagreement between two draws gives one share. More cheap draws or an agreement threshold would trace a curve; that is the next run.
Next
Swaptasks.jsonl for a sample of your own traffic and the verifier for one
that grades it; the routers and the controls do not change. Before you build
a router, run --dry-run with your pair, then live with --limit 200: if
cascade does not beat random on 200 of your questions, a router will not
pay for itself.
References
- Wang et al. 2024. MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark. arXiv:2406.01574.
- Ong et al. 2024. RouteLLM: Learning to Route LLMs with Preference Data. arXiv:2406.18665.
- Chen, Zaharia, Zou 2023. FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance. arXiv:2305.05176.
- Wang et al. 2022. Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv:2203.11171.