Eight steps
Steps 02, 03, 04 and 07 are ours. The rest is the literature. Code paths
are relative to
whileai/simulations/.
Measurement
The held-out tasks are the only number that counts. Everything else on
this page exists to make that number mean something.
Questions
Importance sampling? No. We cover the failure space rather than estimate production. A row carries the policy version, and withlogprobs=True the log-probabilities an off-policy trainer needs to form
the truncated ratio exp(log pi_new - log pi_old) itself [10, 11].
SFT or RL? Both, from the same graded rows. Reward=1 rows for SFT,
pairs for DPO, groups for GRPO, everything for a reward model.
The judge is another LLM. Yes. So it is measured against gold labels,
probed with known hacks, versioned by rubric hash, and drawn from a
different model family than the policy [8].
How do you know training helped? delta_report: paired before and
after on held-out tasks with a bootstrap interval. A must_not_regress
marker whose interval sits below zero fails the run; markers such as
argument grounding catch what pass@1 hides [1].
References
- Lambert, N. Reinforcement Learning from Human Feedback, chapter Evaluation. 2025.
- Chen, M. et al. Evaluating Large Language Models Trained on Code. 2021.
- Yao, S. et al. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. 2024.
- Shao, Z. et al. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. 2024.
- Yu, Q. et al. DAPO: An Open-Source LLM Reinforcement Learning System at Scale. 2025.
- Kuhn, D. R., Wallace, D. R., Gallo, A. M. Software Fault Interactions and Implications for Software Testing. IEEE TSE 30(6), 2004.
- Lehman, J., Stanley, K. O. Abandoning Objectives: Evolution Through the Search for Novelty Alone. Evolutionary Computation 19(2), 2011.
- Zheng, L. et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023.
- Gao, L., Schulman, J., Hilton, J. Scaling Laws for Reward Model Overoptimization. ICML 2023.
- Schulman, J. et al. Proximal Policy Optimization Algorithms. 2017.
- Noukhovitch, M. et al. Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models. ICLR 2025.
What to run next
recipes/01-simulate/bring-your-own-agent
runs these eight steps on a callable of your own, offline and in seconds.