Skip to main content
The scripts are in the repository, not in the installed package. Clone it, then cd recipes/02-measure/model-router before running the commands below. Browse this recipe on GitHub.
Send each request to a cheap model or to the frontier model, and measure whether the router earns its keep. A router is only worth building if it beats sending the same share of traffic to the frontier model at random, so every router here is scored against that control, paired on held-out questions, with a 95% interval. What you will learn: why “the router keeps 90% of quality at 30% of the cost” means nothing without the random control at the same share; that a router reading only the prompt barely beats random on this traffic, while a cascade that reads the cheap model’s own answers does; and how large the gap to a perfect router still is. You need an OpenAI-compatible key (OPENAI_API_KEY; an OpenRouter key works with the default --base-url). --dry-run needs nothing. About 12 minutes and $5.71 live, 15 seconds offline.

Run it

The setup

Traffic. 1,190 MMLU-Pro test questions (Wang et al. 2024 [1]), 85 from each of 14 subjects, frozen in tasks.jsonl by make_tasks.py. They stand in for an internal assistant’s mixed traffic: law, business, health, engineering and so on, ten options each, graded by wai.verify.MultipleChoice with no judge. A third of each subject is held out by a hash of the question id (772 train, 418 held out). Nothing a router learns comes from the held-out third. Models. Gemma 3 12B (open weights, one GPU) and Claude Sonnet 5.5, both through OpenRouter at temperature 0.7. The cheap model answers every question twice. The frontier model answers train once and held-out twice, so every number below is read on a second, independent draw too. Routers. Each sends the top share of held-out questions to the frontier model:
  • subject: the free rule. Escalate the subjects where the frontier model gained most on train.
  • text: logistic regression on hashed word n-grams of the question, predicting “the cheap model misses” from its two train draws (the win-label router of RouteLLM [2], in numpy).
  • cascade: answer with the cheap model twice; escalate when the two answers disagree (FrugalGPT [3], with self-consistency [4] as the confidence). It pays for two cheap calls on every question, and its share is whatever the disagreement rate turns out to be.
  • oracle: escalate exactly the questions where the frontier model is right and the cheap one is wrong. The ceiling, not a router.
Controls. vs random compares each router with random routing at the same frontier share, written as 40 rollouts per question so the control is exact rather than one lucky draw. vs all frontier compares it with the baseline it is trying to replace. Both come from wai.compare, paired by question.

Result

Run 2026-10-01, every call made in that run, nothing reused. Held out, draw 1: python run.py prints all shares (0.2, 0.4, 0.6) and both draws; results.json has every number. The cascade is the only router that beats random on both draws: +5.4 points [+2.4, +8.4] on draw 1, +4.5 [+1.4, +7.5] on draw 2. It is the only router that reads the cheap model’s answers rather than the question. The prompt-only routers are about the same as random. The text router’s interval crosses zero at every share on both draws, and its AUC for “the cheap model will miss” is 0.588. The subject rule clears random on draw 1 (+4.6 at 40%) and does not replicate on draw 2 (+2.5 [−0.1, +5.4]). What a question looks like says little about whether a 12B model will get it wrong. The room is real. The oracle sends 33% of traffic to the frontier model and beats all-frontier by 6.9 points [+4.8, +9.3] at 0.39x the cost, because on about 7% of questions the cheap model is right and the frontier model is wrong. A router that found the third of questions that need the frontier model would win on both quality and cost. None of the three here comes close. No router keeps quality within 2 points of all-frontier below 95% frontier share. Read that as the honest answer to “how much traffic can we move off the frontier model for free” on this traffic: almost none, with these routers.

Limits

  • MMLU-Pro is a stand-in for your traffic, and a public one. Either model may have seen the questions. Real enterprise traffic has more easy questions, which helps every router.
  • Costs are list prices times tokens used. Escalated questions cost more than average (they are longer and get longer answers), so a 0.26 share costs 0.40x, not 0.26x.
  • One train split, one router fit. The second draw replicates the evaluation, not the router’s training.
  • The cascade has one operating point. Disagreement between two draws gives one share. More cheap draws or an agreement threshold would trace a curve; that is the next run.

Next

Swap tasks.jsonl for a sample of your own traffic and the verifier for one that grades it; the routers and the controls do not change. Before you build a router, run --dry-run with your pair, then live with --limit 200: if cascade does not beat random on 200 of your questions, a router will not pay for itself.

References

  1. Wang et al. 2024. MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark. arXiv:2406.01574.
  2. Ong et al. 2024. RouteLLM: Learning to Route LLMs with Preference Data. arXiv:2406.18665.
  3. Chen, Zaharia, Zou 2023. FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance. arXiv:2305.05176.
  4. Wang et al. 2022. Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv:2203.11171.
Last modified on October 1, 2026