- Simulate. Sample a group of rollouts per task. Passes differ: one looked the order up first, one refunded twice, one invented an id that happened to exist.
- Grade. A program says pass or fail. Every pass scores 1.0, so the policy learns to pass by any means and rollouts grow.
- Measure. Before training, score a sample of passes with the grader. If the scores do not vary, stop: a constant is erased by group normalization.
- Select. Give the grader one whole group. It ranks the passes and names the hacks; ranks become quality factors.
- Train. Move the positive advantage from lower- to higher-quality passes with its total conserved, then train as before.
whileai writes no loss for this.
It carries the grader and the report’s knobs as one object, hands TRL a
reward function whose group normalization reproduces the redistributed
advantages, and refuses to run a grader that cannot tell the passes apart.
What the research says
- Two places to grade. Groupwise Reward Synthesis (GRS, section 4.3.1)
is offline: a grader reads a group of rollouts plus the task and writes
task-specific rubrics (solution criteria and behavior criteria); during
training each pass is scored against them and the reward is the product
R_i = R_test_i * S_sol_i * S_beh_i(equation 2). Groupwise Advantage Redistribution (GAR, section 4.3.2) is online: for each mixed-outcome group the grader inspects every rollout jointly, ranks the passes on approach, precision, minimality, side effects and craftsmanship, zeroes a confirmed hack (an external or leaked answer) and treats it as a failure before the group statistics are recomputed. The report uses GRS on a subset of high-pass-rate tasks and GAR on everything else [1]. - The redistribution conserves mass. With
A_i = R_i - mean(R)andPthe passes, quality factorsf_iin (0, 1] first downweight the lower-quality passes, then a common factorlambda = sum_P A_j / sum_P f_j A_jputs the removed mass back among the passes:A'_i = lambda f_i A_ifor a pass,A_iotherwise (equation 3). Failures are untouched,sum_P A'_i = sum_P A_i, and the relative weights are the grader’s. The report caps lambda against runaway amplification and then subtracts the group mean again so the group has zero mean; unusable grader output falls back to the original advantages [1]. - A grader with no spread is a null result. On 2026-09-21 GRS was run on single-turn text-to-SQL with Qwen3-4B: the rubric grader gave a mean of 0.93 to 1,472 passing replies, the multiplier was a constant, GRPO’s group normalization erased it and the arm matched the plain-reward arm (-1.7 points, 95% -5.2 to +1.7) at $14. Two rules follow. Check the grader’s spread on a sample of passes before the GPU is spent. And put the method on multi-turn agent tasks where passes differ (an extra tool call, a skipped verification, an invented id), not on one-line answers.
The calls
The grader is your judge call. In advantage mode it receives one group as a list of{"index", "prompt", "completion", "reward", "passed"} and
returns a best-first ranking over the passing indices (an inner list is
a tie), or factors in (0, 1] by index, plus an optional hacks list.
The scripted grader below ranks by how many tool calls the reply made,
fewest first, and calls a pass that names an order id no tool returned a
hack. A real one is a model reading the group.
scenario_id, reward and final_text) and let it grade the
groups the way training would. No spread is a ValueError that says what
to change; strict=False returns the report instead.
GRPOTrainer calls a reward function on a batch
of num_generations consecutive completions per prompt and forms the
advantage as (r - mean_group) / std_group. trl_reward scores the batch
with your verifier, grades each mixed group, and returns A'_i + mean(R),
so TRL’s mean subtraction gives equation 3 exactly. The division by the
group’s standard deviation is TRL’s and is the one limit: it rescales the
redistributed group by its new spread, so the ranking and the relative
weights hold but the total moved mass matches the report only under
scale_rewards="none" (Dr. GRPO [2]). Set it the same way in both arms of
a comparison.
reward_funcs=[reward_func] to GRPOTrainer with use_vllm=True,
vllm_mode="colocate" and scale_rewards="none". gar.stats prints
groups seen, grader calls and failures, hacks zeroed and the spread of the
factors, so the run’s grader is a measurement too.
Reward mode is the same object with mode="reward": the grader receives
one passing rollout and its rubric and returns
{"solution": s, "behavior": b}; the reward becomes
R_test * max(floor, s) * max(floor, b) with floor 0 as in equation 2.
rubrics is a mapping from task_id to rubric, or a callable that writes
one from the first group seen for a task and is cached.
The defaults, and where they come from
Every one is a named constant in
whileai/simulations/defaults.py
(GROUPWISE_*, SPREAD_*) and a field on the object.
References
- Xiaomi MiMo team. MiMo-V2.6 Technical Report, 2026-09-21, section 4.3 “Groupwise Agentic Grading”: 4.3.1 Groupwise Reward Synthesis, 4.3.2 Groupwise Advantage Redistribution, Figure 8.
- Liu et al. Understanding R1-Zero-Like Training: A Critical Perspective (Dr. GRPO), 2025, arXiv:2503.20783.
- Lambert. Reinforcement Learning from Human Feedback, 2025, arXiv:2504.12501, chapter “Over-Optimization”.