The scripts are in the repository, not in the installed package. Clone it,
then
cd recipes/01-simulate/swarm-rescue before running the commands below. Browse this recipe on GitHub.WHILEAI_API_KEY for the hosted Qwen3-4B (wai login), or any
OpenAI-compatible endpoint via --base-url; pyarrow for the dataset.
Offline with --dry-run. Takes: seconds offline; about three and a half
hours for every task and two noise re-runs on the hosted model.
The problem
GRPO learns from the spread inside a group of rollouts. On a task where all 8 fail, the advantage is zero and dynamic sampling drops the prompt [1]. The SDK does the same:simulate(mode="rl") fills a group only when its first
two rollouts disagree. Those tasks are the hardest ones in the set, the
ones a training run most needs, and “sample more” rarely helps: the
policy puts almost no mass on the answer, so 24 more independent draws are
24 more misses.
A particle swarm [2] changes what the extra samples are conditioned on.
Each particle keeps its best attempt so far, sees the test it failed, and
sees a neighbour’s best attempt. The LLM is the velocity update: given
its own best and a neighbour’s, it writes the next program. Topology is
the variable. In a ring each particle sees two neighbours, so good ideas
spread slowly and the swarm stays diverse; in a star everyone sees the
global best, which is faster and collapses the swarm onto one idea [3].
The recipe
-
Tasks:
deepmind/code_conteststest and valid splits [4], 275 Codeforces-style problems after dropping the interactive ones, median rating 1900. The visible tests are the public tests plus 8 generated ones; the hidden tests are the private tests plus 32 more generated ones. Pass = the program’s output, split on whitespace, matches on every test within 6 seconds, numbers within 1e-6. - Base: 8 rollouts per task from the hosted Qwen3-4B, thinking off, temperature 1, 2,048 tokens. A task where all 8 fail every test is an all-fail task. The script prints the per-task histogram, not the mean, because the mean hides the shape.
-
Every all-fail task gets 24 more samples four ways:
- resample: 24 fresh independent samples.
- solo: 8 particles, 3 rounds. Round 0 is 8 fresh samples. In each later round a particle sees its own best attempt (most visible tests passed; ties go to the newer one) and the first visible test it failed, with the input, the expected output and what it printed, and writes a new program. No sharing.
- ring: solo, plus the best attempt of the particle’s two ring neighbours, shown beside its own.
- star: solo, plus the best attempt of the whole swarm.
-
Measure: rescue rate per arm, the share of all-fail tasks with at least
one rescue, with a bootstrap interval over tasks (
wai.pass_at). Each swarm arm against resample is a paired delta over the same tasks with a sign-flip p-value (wai.compare_runs).--noise-runs 2re-runs the resample arm with new seeds and reports the band a delta has to clear (wai.eval_variance). -
Export: every rescued program goes to
out/rescued.jsonlas a bare row (the problem in, the passing program out) with no swarm context in it. That file is what a later SFT or distillation run trains on; the swarm prompts never leak into it.
Result
Run 2026-09-21 on the hosted Qwen3-4B. 275 tasks, 223 all-fail at 8 rollouts (81%). Per-task passes out of 8: 223 tasks at 0, then 17, 14, 4, 3, 1, 3, 4, 6.
Flat. Both swarms rescued 22 tasks to resampling’s 16, and the gap sits
inside its interval. The noise band says the same thing louder: the
resample arm re-run with two more seeds rescued 20 and 24 tasks (9.0% and
10.8%), so its run-to-run spread is 1.8 points and a delta under 10.9
points is the eval re-running. The swarms’ 9.9% is where resampling lands
on an average seed.
Replicated. A second seed of the whole experiment (
--seed 1,
228 all-fail tasks) put every arm between 10.1% and 11.4%, resample re-runs
at 11.4, 10.1 and 11.8%, and a band of 5.6 points. Ring came in 1.3
points behind resampling and star level with it. The seed-0 lead for the
swarms was the draw.
What did move is the training set. Across the six runs (four arms and two
re-runs, 144 extra samples a task), 43 of the 223 all-fail tasks got at
least one program passing every hidden test, 19% of the prompts a single
8-rollout group would have dropped; seed 1 rescued 49 of 228 the same way. The four arms alone wrote 112
distinct passing programs on 37 tasks into
out/rescued.jsonl.
Raising the floor
The flat result had one candidate cause: fitness was zero almost everywhere, so a neighbour’s best program carried no information. Two more configurations tested that, same recipe, same measurement. Thinking mode on the 4B is out. On a hard task 32 of 40 replies ran past 6,144 tokens still thinking and never wrote a program. Qwen3.8-27B, thinking off, is in (--base-url, --max-tokens 8192,
--no-comments: with thinking off it reasons inside the code block as
comments without end, so the prompt gains one line and the cap doubles,
both recorded in results-27b.json). It fails all 8 rollouts on 107 of
275 tasks. On 57 of those at least one base rollout passes a visible test,
the near-miss band where fitness is graded (--band near-miss): the best
base attempt passes a third of the visible tests on average, and half of
them on one task in five.
Flat again. The resample re-run landed at 29.8%, and 29 of the 57 tasks
were rescued by at least one of the five runs;
rescued-27b.jsonl holds
108 passing programs on 28 tasks. Cost: about 8,000 calls and 40M tokens
over four hours on the SGLang endpoint, with 22% of arm replies still
at the 8,192 cap.
Learned
- Feedback alone did nothing here. Solo refinement, the swarm with the social term off, matched resampling exactly. On the unrescued tasks the best program passes 8% of the visible tests on average, so the fitness signal the swarm moves on is almost always zero, and a particle shown a neighbour’s zero-fitness program has no direction to move in. A swarm needs a graded fitness; pass or fail on a task the model cannot touch is not one.
- The rescues that did happen came in the refinement rounds for the
swarms (17 of 22 for ring, 16 of 22 for star) and in round 0 for
resampling (8 of 16), which is the pattern the design predicts. It is
the size that is missing:
wai.holdout_sizesays a +5 point gap at a 7% base needs 547 paired tasks and a +3 point gap needs 1,382. 223 was never going to resolve a gap this small. - A real gradient did not change the answer. On the 27B’s near-miss band the swarm had a third of the visible tests to climb and still tied 24 independent draws, with most rescues in round 0 for every arm. Three configurations, three flat results: PSO over rollouts, as a way to rescue the prompts GRPO drops, is closed. What a hard prompt needs is more samples or a stronger model, not a smarter way to condition them.
- The dataset angle survives the flat result. On tasks the model gets right 1 time in 30 or less, how you structure the extra samples did not matter; that you spend them did. Six runs of 24 turned 43 zero-gradient prompts into learnable ones. The next recipe trains on those rows (rejection-sampled SFT or on-policy distillation from the bare prompt, since the swarm rows are off-policy for it) against plain GRPO at matched rollouts.
results.json and
results-seed1.json in this directory are the two runs’ reports;
rescued-seed1.jsonl is seed 1’s 127 passing programs on 42 tasks, the
rows a training run starts from. python run.py --reuse reprints a report
from out/.
References
- Yu, Q. et al. DAPO: An Open-Source LLM Reinforcement Learning System at Scale. arXiv:2503.14476, 2025.
- Kennedy, J., Eberhart, R. Particle Swarm Optimization. IEEE ICNN, 1995.
- Kennedy, J., Mendes, R. Population Structure and Particle Swarm Performance. IEEE CEC, 2002.
- Li, Y. et al. Competition-Level Code Generation with AlphaCode. arXiv:2203.07814, 2022.
- Feng, S. et al. Model Swarms: Collaborative Search to Adapt LLM Experts via Swarm Intelligence. arXiv:2410.11163, 2024. PSO over LoRA weights; this recipe runs PSO over programs instead.
- Li, J. et al. QuestA: Expanding Reasoning Capabilities of LLMs via Question Augmentation. arXiv:2507.13266, 2025. Rescues the same zero-advantage prompts with partial reference solutions as hints; the swarm makes its hints from its own population.