Skip to main content
The scripts are in the repository, not in the installed package. Clone it, then cd recipes/04-train/model-router before running the commands below. Browse this recipe on GitHub.
A model router reads a question and picks one model from a pool, trading accuracy against cost with one knob. This recipe trains three routers from the 2025-2026 research on twelve frontier models and scores them on held-out questions against the best single model, a random pick, the perfect router, and OpenRouter’s auto router. It then asks an untrained decision model, TypeSafe’s Jev, to route the same questions. What you will learn: how a router is trained (it is a table lookup more than a neural network); why the cluster-based router from Avengers-Pro [1] is the one to start from; what the knob does to cost and accuracy; how far every router still sits from the perfect one; and why a model that only names the kind of question can match them. You need numpy, fastembed and huggingface_hub, and no API key: the answers were already collected by LLMRouterBench [2]. Only the optional Jev section needs a TypeSafe key. --dry-run needs only numpy. prepare.py takes about 10 minutes on a laptop CPU once (it downloads 1.3 GB); run.py takes about 2.

Run it

How a router is trained

The data is a table. For every question, LLMRouterBench [2] ran every model, graded the answer and recorded what it cost. prepare.py keeps the twelve flagship models (GPT-5, GPT-5 Chat, Gemini 2.5 Pro and Flash, Claude Sonnet 4, DeepSeek R1 and V3, two Qwen3-235B, Kimi K2, GLM-4.6, Intern-S1) and the twelve datasets all of them answered: AIME, ARC-AGI, three ArenaHard splits, GPQA, HLE, LiveCodeBench, LiveMathBench, MMLU-Pro, SimpleQA and SWE-bench Verified. Each dataset is capped at 1,000 questions, which leaves 5,529. Every question is embedded with a 33M-parameter encoder on CPU; the benchmark found the choice of encoder matters little [2]. Training is reading the table by neighbourhood. All three routers answer one question: on questions like this one, how did each model do, and what did it cost?
  • knn [3]: find the 50 most similar training questions and average each model’s score and cost over them.
  • avengers-pro [1]: cluster the training questions into 25 groups. In each group, score every model on accuracy (normalised across models) and on cost. A new question reads its 3 nearest groups. The knob weighs accuracy against cost. The settings are the reference implementation’s in LLMRouterBench.
  • linear [4]: one ridge regression per model, from embedding to score: the parametric baseline, which predicts each model separately and can flip rankings on small errors [5].
The knob. Each router sweeps a cost weight from 0 (accuracy only) to 1 (cost only). Questions are split 80/20 by a hash of their id (the benchmark’s own ratio). The 80% train part is split again, and the knob is chosen on a validation slice for three operating points:
  • best accuracy;
  • the cheapest knob that matches the best single model’s accuracy;
  • the most accurate knob that costs no more than OpenRouter’s auto router.
The router is then refit on all of train and read once on the 1,061 held-out questions. Every comparison is paired by question through wai.compare. ArenaHard ties (score 0.5) are written as one pass and one fail out of two.

Result

Held out, 1,061 questions. The routers ran on 2026-10-01 over answers LLMRouterBench collected in late 2025: no model was called for this recipe. Accuracy against cost per 1,000 questions for twelve models, the Avengers-Pro and kNN routers, and the perfect router The router keeps the best model’s accuracy at 38% of its cost. Avengers-Pro at the matched point is about the same as Gemini 2.5 Pro (+0.1 points, interval [−2.7, +2.9]) for $34 per thousand questions instead of $91. Against GPT-5, which is about as accurate and half Gemini’s price, the saving is 24%. That is close to what Avengers-Pro’s authors report against GPT-5 (27% [1]). It sends 55% of questions to GPT-5, 23% to Qwen3-235B and the rest to Kimi K2 and the two reasoning models. No router is clearly better than the best model, at any price. The best accuracy any router reached is +1.8 points [−0.8, +4.3]. The benchmark reports up to +4% [2]. Against OpenRouter’s auto router, on the 915 held-out questions the benchmark ran it on: Accuracy of OpenRouter's auto router, Avengers-Pro and kNN at or under OpenRouter's cost Avengers-Pro is 5.6 points more accurate at 74% of OpenRouter’s cost. Read this as a result about the auto router the benchmark called in late 2025. OpenRouter has since replaced it with one that picks by task type and community spend [6]. Every router is stuck about 19 points below the perfect one. The oracle reaches 0.822 at a tenth of Gemini’s cost. All three routers land within 2 points of each other, which is the routing plateau [7]: routers that read only the question learn which model is good at which kind of question, and miss the questions only one or two models get right.

A Jev router, untrained

Can a model that was never trained on this table route as well as one that was? jev.py sends each val and held-out question to TypeSafe’s Jev [9], a decision model that answers typed questions with a probability per answer and writes no text. One request asks two things:
  • jev-task: which of the twelve kinds of question is this? Each kind is one sentence (“A short factual question with one specific answer”). run.py averages each model’s training accuracy and cost over the kinds Jev thinks the question could be. This is a task-type router, like OpenRouter’s [6].
  • jev-pick: which model will answer correctly? Each model is listed with its training accuracy on every kind of question.
Jev named the kind of question correctly 75.7% of the time. That was enough. jev-task matched Avengers-Pro and beat it at low budgets. Both knobs were picked on val at the same budget and compared on the same held-out questions: At $5, Jev is 1.6 points more accurate for less money. From $20 up it spends more than Avengers-Pro on held-out questions, so the comparison is no longer at equal cost. At the top budgets the two tie. Asking Jev for the model directly (jev-pick) did worse than asking it for the kind of question: 0.617 against 0.636 at the matched-accuracy point. Jev is better at reading a question than at reading a table of percentages. Two caveats. The twelve kinds are the twelve source datasets, which is the best case for a task classifier: real traffic does not arrive sorted into twelve benchmarks. And Jev’s median reply took 0.27 s with eight requests in flight. That is fast enough to update a suggestion after the user stops typing, but not on every keystroke. The embedding routers answer in milliseconds on a CPU. The result fits the plateau [7]. Jev reads only the question, like the other routers, so it lands on the same line. It gets there without training.

Limits

  • One benchmark, answers from late 2025. Prices and models have moved, and so would the routing.
  • The held-out set and the training set come from the same twelve datasets. A router on traffic unlike any of them would do worse.
  • One train/test split. K-means is the only random step: over five seeds Avengers-Pro at the matched point ranges 0.615 to 0.623 in accuracy and $32.82 to $34.41 per 1k. The point closest to OpenRouter’s cost varies more ($21.60 to $32.66), because two knob settings sit close on val.
  • Costs are what the benchmark recorded, including reasoning tokens. A router that picks a reasoning model pays for its thinking.

Next

Part 2, decision-router, works out what a decision model like Jev is from the outside and trains small models of its shape on this table. The plateau says the next gain comes from reading more than the question: the start of a model’s answer [7], or a small open model’s internal activations as it reads the prompt [8]. To route your own traffic, replace the table with your questions, each model’s graded answer and its cost (recipes/02-measure/model-router collects exactly that for two models), and keep run.py.

References

  1. Zhang et al. 2025. Beyond GPT-5: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing. arXiv:2508.12631.
  2. Li et al. 2026. LLMRouterBench: A Massive Benchmark and Unified Framework for LLM Routing. Findings of ACL 2026. arXiv:2601.07206.
  3. Li 2025. Rethinking Predictive Modeling for LLM Routing: When Simple kNN Beats Complex Learned Routers. arXiv:2505.12601.
  4. Hu et al. 2024. RouterBench: A Benchmark for Multi-LLM Routing System. arXiv:2403.12031.
  5. Lai and Ye 2026. When Routing Collapses: On the Degenerate Convergence of LLM Routers. arXiv:2602.03478.
  6. OpenRouter. Auto Router. openrouter.ai/docs/guides/routing/routers/auto-router, read 2026-10-01.
  7. Lu et al. 2026. The Routing Plateau: Understanding the Accuracy Limits of LLM Routers. arXiv:2606.07587.
  8. Varshney et al. 2026. LLM Router: Rethinking Routing with Prefill Activations. arXiv:2603.20895.
  9. TypeSafe AI. Jev, System One API. api.typesafe.ai/v1/systemone, model jev-1.13.0, called 2026-10-01.
Last modified on October 2, 2026