The scripts are in the repository, not in the installed package. Clone it,
then
cd recipes/04-train/decision-router before running the commands below. Browse this recipe on GitHub.model-router) an untrained decision model, TypeSafe’s
Jev, routed twelve frontier models as well as the routers trained for it.
This recipe works out what Jev is from the outside, then trains three small
models of that shape on the same table and puts them, Jev and part 1’s best
router through one test: the same budgets, the same 1,061 held-out
questions, every pair compared question by question.
What you will learn: what is known about how Jev is built and how the
evidence was gathered; how a decision model reads a question once and
scores every candidate in one forward pass, with no text generated; why it
is trained with log loss and then a single temperature; and whether a
0.6B model you train in minutes matches Jev on routing. You need part 1’s
out/ (run its prepare.py, and its jev.py for the Jev rows), a Modal
token for training, and a TypeSafe key only for probe_jev.py.
--dry-run needs only numpy.
Run it
What Jev is
TypeSafe has published no paper, weights or parameter count. What it says [1, 2]: Jev is a transformer with a new architecture and a “parallel sampler”, trained on synthetic data only with a method it calls reinforcement learning for calibrated decisions (RLCD). It takes a block of state and typed questions (pick one option, score on a scale, or say whether a statement is true) and returns probabilities, never text. One limit gives the shape away: a request may carry 32k tokens of “state plus the longest question” [2], so each question is its own sequence over a shared state. Two outside studies fill in the rest. Hume [3] made more than 10,000 requests. Jev’s tokenizer matched Qwen’s on 348 of 415 probes, with some merges changed. Theoutput_tokens it bills is counted from the reply, not
generated. Latency grows with the state and barely with the number of
questions. Adding an irrelevant option shifted the odds between two others
(−0.28 log-odds [−0.36, −0.19]), so the options are read together, not
scored one by one. Deußer et al. [4] ran Jev on 37 datasets and found its
choice probabilities well calibrated.
probe_jev.py repeats the checks that matter for training a copy
(probe.json holds one run, 2026-10-01, jev-1.13.0):
The best-supported reading: a causal transformer, likely Qwen-derived,
reads the state once and shares it; each question is a separate branch over
it; a head reads a probability for every option from that branch, with the
options listed in the input; and the training rewards honest probabilities
with a proper scoring rule, the property RLCR proves for a Brier reward [8].
The size (Hume estimates about 10B active parameters) and whether the
backbone is a mixture of experts are guesses.
Three models of that shape
train_modal.py trains both on part 1’s split (3,305 questions to fit,
1,163 to pick the epoch, the temperature and the knob, 1,061 held out). The
target for each question is twelve numbers: whether each model got it right
(ArenaHard ties count 0.5).
pointer, the Jev shape. Qwen3-0.6B [9] with a LoRA adapter [10] readsQuestion: ..., then the twelve model names, each followed by a reserved marker token that question text cannot forge (Hume found Jev’s option boundaries cannot be forged either [3]). A linear head reads the hidden state at every marker: one logit per model. Later names attend to earlier ones, so the list is read together. The list is shuffled on every training pass so no position means anything, and val and test average four shuffled orders.encoder, the fast baseline the open copies use [11]. ModernBERT-base [12], 150M parameters, reads the question; a fixed twelve-way head scores the models. It cannot take a model it was not trained on, which the pointer can.pointer-kind, the pointer plus the kind of question. Jev names the kind (part 1’sjev.py --allasks it about every question, training ones included), and part 1’s table turns that into each model’s expected accuracy, asjev-taskdoes. That prior is added to the pointer’s logits with a learned weight; the head starts at zero, so training starts from the prior and learns a correction from the question text. A training question is left out of its own kind’s average, so its label never feeds its own input.
Result
Held out, 1,061 questions. Trained on 2026-10-01: nine runs on Modal L40S, about an hour and about $8 in all. Jev’s rows are part 1’s, from the same day.
Each cell is accuracy, with the held-out cost per 1,000 questions in
brackets. The knob is picked on val, so a router can spend somewhat more or
less than the budget on held-out questions.
Our pointer model ties part 1’s best router and loses to Jev at low
budgets. Paired, pointer minus Avengers-Pro is −1.4 points [−3.0, +0.2] at
$5 and +0.4 [−1.4, +2.0] at $35: no difference we can see. Pointer minus Jev
is −3.0 [−4.7, −1.5] at $5 and −2.2 [−4.2, −0.3] at $10. From $20 up the
interval covers zero. The other two seeds land within 1.5 points of the
reported one at every budget. The encoder does the same or slightly worse.
Our models estimate each model’s chance better, and it does not turn into
better routing. Over every (question, model) pair on held out, the
pointer’s probability of “this model gets it right” has a Brier score of
0.162 (seeds 0.161 to 0.165) and a calibration error of 0.013 to 0.037. kNN’s
neighbour averages, the input to part 1’s routers, score 0.190 and 0.025. A
better estimate per model still picks about the same model, because the
twelve flagship models get mostly the same questions right. What decides a
pick is which model is better on this question, and the question alone
barely says. That is the routing plateau again [14].
Jev’s edge is that it recognises the benchmark. The control row routes
with each question’s true dataset and part 1’s table: 0.581 at $5 and 0.585
at $10, against Jev’s 0.576 and 0.583. Jev names the dataset right 75.7% of
the time, and that accounts for its lead. On real traffic there are no
twelve tidy benchmarks to recognise, so this is the part of part 1’s result
least likely to carry over. A 3,305-question training set did not teach our
models to separate the twelve sources as sharply as a model that reads them
in plain English.
Given the kind, our model matches Jev, and the question text adds little
on top.
pointer-kind minus Jev is −0.6 points [−1.5, +0.2] at $5,
−0.1 [−1.4, +1.3] at $10 and +0.3 [−0.4, +1.0] at $50: a tie at every
budget, where the plain pointer trailed by 3.0 at $5. Against Avengers-Pro
it is ahead at every budget, and the interval clears zero at $10 (+1.7
[+0.1, +3.3]). Its other two seeds score up to 1.5 points lower. But
matching jev-task is what the prior alone already does; the pointer’s
correction from the question text is within a point of nothing at every
budget. And it needs Jev’s answer at inference. On your own traffic the
kind would come from a classifier over your own categories, and how well
that works depends on how well your categories separate the models.
Speed. One question at a time on an L40S: the encoder takes 14 to 20 ms,
the pointer 30 to 60 ms (twelve markers, one order; pointer-kind adds a
Jev call or your own classifier). Jev took 0.27 s median
over the network with eight requests in flight. Either of ours is fast
enough to update the pick as someone types; Jev is fast enough to update it
when they pause.
Limits
- One benchmark and one split, as in part 1. Three seeds per model give the training spread; they say nothing about other traffic.
- The trained models never see costs or the knob. They estimate correctness only, and part 1’s knob trades it against each model’s mean cost. A model that predicts cost per question too (reasoning models spend very different amounts on different questions) might pick better.
- Small, untuned models. One size each, four epochs, LoRA rank 16, the learning rates are conventions. A larger base, or a pointer started from a model that already reads lists of options, might add more on top of the prior; we did not test it.
- Jev’s architecture is inferred, not known. The pointer copies the shape the evidence supports, not Jev itself.
Next
pointer-kind shows where the gain was: the kind of question carries the
signal, and the question text adds little beyond it. On your own traffic,
start with the categories you already have and a table of how each model
does on each. A reply to part 1 suggested the next input: let people
overrule the pick, and train on the overrules. They are a label no
question-only router has [14].
References
- TypeSafe AI. Introducing System One Models & Jev. typesafe.ai/blog/introducing-system-one-models-and-jev, 2026-09-15.
- TypeSafe AI. Models (model card, jev-1.13.0). docs.typesafe.ai/models, read 2026-10-01.
- Hume, A. Jev’s Architecture Unmasked. archerhume.com/posts/jevs-architecture-unmasked, 2026-09-17.
- Deußer, Sparrenberg and Sifa 2026. Evaluating and Benchmarking the System One Model Jev. arXiv:2609.37647.
- Zheng et al. 2024. Large Language Models Are Not Robust Multiple Choice Selectors. ICLR 2024. arXiv:2309.03882.
- TypeSafe AI. Jev 1.13 jaggedness. docs.typesafe.ai/model-jaggedness/jev-1.13, read 2026-10-01.
- Juravsky et al. 2024. Hydragen: High-Throughput LLM Inference with Shared Prefixes. ICML 2024. arXiv:2402.05099.
- Damani et al. 2025. Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty. ICLR 2026. arXiv:2507.16806.
- Qwen Team 2025. Qwen3 Technical Report. arXiv:2505.09388.
- Hu et al. 2021. LoRA: Low-Rank Adaptation of Large Language Models. arXiv:2106.09685.
- wfzyx. Von: the open-source System One decision model (ModernBERT, 395M). github.com/wfzyx/von, read 2026-10-01.
- Warner et al. 2024. Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference (ModernBERT). arXiv:2412.13663.
- Guo et al. 2017. On Calibration of Modern Neural Networks. ICML 2017. arXiv:1706.04599.
- Lu et al. 2026. The Routing Plateau: Understanding the Accuracy Limits of LLM Routers. arXiv:2606.07587.