The scripts are in the repository, not in the installed package. Clone it,
then
cd recipes/community/deepagents-review-four-arms before running the commands below. Browse this recipe on GitHub.wai.simulations.optimize keeps. Nothing else differs. Both models are scored on SWE-bench Verified
patch review, a benchmark nobody here built, and on 84 repositories the traces never touched.
What you will learn: that which traces you train on moves a fine-tune as much as the trainer
does, how to hold a comparison to one variable (same trainer, same server, same sampling), and
why a single base run is not a baseline. You need OPENROUTER_API_KEY and LANGSMITH_API_KEY
for traces, Docker for smithtune (it imports fcntl, so Linux or macOS), and a Modal account for
training and serving. python run.py --dry-run needs none of it.
Result
SWE-bench Verified patch review, 250 reviews on 125 issues, 50/50 approve/reject, labels from the official evaluation. Every arm on one vLLM server, thinking off, same sampling:
While’s pick against smithtune’s, paired by review: +10.8 points [+4.8, +16.8] and 5.3 fewer
calls [−6.7, −4.0], on 65% fewer training tokens per epoch (3.6M against 10.2M). The noise floor
from three base runs (
wai.eval_variance, t at 2 df) is 7.8 points; the gain clears it.
On the 84 held-out repositories (246 reviews): 61.4% against 63.8%, a tie (−2.4 [−7.3, +2.4]),
and 3.1 fewer calls [−4.3, −2.0] for While’s pick. All numbers are in results.json
(python make_results.py).
The write-up with charts: Fine-tune a Deep Agents reviewer from its LangSmith traces.
How it works
data.pybuilds the review tasks fromnebius/SWE-agent-trajectoriesandnebius/SWE-bench-extra(pinned revisions): for each issue with a passing and a failing SWE-agent patch, one of each. Splits are by repository.wai.decontaminatechecks train and dev against the holdout (0 of 1,306 dropped, every rule ran). Four snapshots deleted from GitHub are listed inunavailable.jsonand left out everywhere.collect.pyruns the stock agent (agent.py:create_deep_agent, read-only filesystem over a tarball of the base commit) and writes LangSmith feedbackcorrectnessandmodel_callson every root run. 400 training reviews: 67% correct.- Without While:
smithtune dataset pullwith the correctness filter, thensmithtune dataset triagewithrubric.md. The council kept 258 of 260. - With While:
pick.pyrunswai.simulations.optimize(mode="sft")with reward 1 only for a correct review within 8 model calls, and writes feedbackwhile_pick = 1; smithtune pulls it with that filter. 173 kept. export_sft.pyruns inside the smithtune image and calls smithtune’sprepare_sft_rows(reasoning omitted, its default),split_rows(80/10/10) andrender_row_tokens(the Baseten renderer: one datum per assistant turn, history at weight 0).smithtune prepareitself stops without a Baseten Loops or Fireworks account; this skips only that preflight.train_modal.pytrains LoRA on one H200 with smithtune’s Baseten defaults: rank 8, batch 32, learning rate 1e-4, seed 42, up to 5 epochs, stop when validation loss does not improve, keep the best epoch. Both arms stopped after epoch 3 and kept epoch 2.serve_modal.pyserves the base and both adapters from one vLLM 0.26 server with the token-level Qwen3.5 tool parser (qwen35_tool_parser.py) andenable_thinking=falsefor every request.public_bench.pybuilds the SWE-bench Verified set from six public leaderboard submissions (s3://swe-bench-submissions,patch.diffandreport.json), none of whose twelve repositories is in the training traces.report.pyandmake_results.pydo the paired comparisons.
Run it
Caveats
- One training seed per arm. The public-benchmark gain clears the three-run noise floor; a second seed per arm is the next run.
- The two training sets differ in size as well as in which traces they hold (173 against 258). A random 173 from the council’s set would separate the pick from the size.
- smithtune omits reasoning from training rows by default, so both tuned models are non-thinking models and are scored against the base with thinking off. With thinking on, through OpenRouter, the base scored 70.3%, 75.1% and 74.6% on the held-out repositories; that run on SWE-bench Verified is still to do. Re-running the same base once moved it 4.5 points with an interval that excluded zero, which is why every comparison here carries the three-run floor.
- Training ran on Modal, not Baseten Loops (not enabled for our workspace), with smithtune’s renderer, split and schedule.
Next
python collect.py --arm base --split public through OpenRouter for the thinking-on base, then a
second seed per SFT arm and the size-matched random control. Traces keep flowing to LangSmith, so
the same pick.py picks the next round.