Skip to main content
Cover the situations. Run them in a sandbox that can fail. Grade with a checked judge. Eight steps from an agent definition to a proven delta; how the simulator thinks and why is How it works. The one-page PDF is at while.ai/while-simulation-engine.pdf. The eight steps, Axes to Delta, with Delta feeding the next run The eight steps, Axes to Delta, with Delta feeding the next run

Eight steps

Steps 02, 03, 04 and 07 are ours. The rest is the literature. Code paths are relative to whileai/simulations/.

Measurement

The held-out tasks are the only number that counts. Everything else on this page exists to make that number mean something.

Questions

Importance sampling? No. We cover the failure space rather than estimate production. A row carries the policy version, and with logprobs=True the log-probabilities an off-policy trainer needs to form the truncated ratio exp(log pi_new - log pi_old) itself [10, 11]. SFT or RL? Both, from the same graded rows. Reward=1 rows for SFT, pairs for DPO, groups for GRPO, everything for a reward model. The judge is another LLM. Yes. So it is measured against gold labels, probed with known hacks, versioned by rubric hash, and drawn from a different model family than the policy [8]. How do you know training helped? delta_report: paired before and after on held-out tasks with a bootstrap interval. A must_not_regress marker whose interval sits below zero fails the run; markers such as argument grounding catch what pass@1 hides [1].

References

  1. Lambert, N. Reinforcement Learning from Human Feedback, chapter Evaluation. 2025.
  2. Chen, M. et al. Evaluating Large Language Models Trained on Code. 2021.
  3. Yao, S. et al. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. 2024.
  4. Shao, Z. et al. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. 2024.
  5. Yu, Q. et al. DAPO: An Open-Source LLM Reinforcement Learning System at Scale. 2025.
  6. Kuhn, D. R., Wallace, D. R., Gallo, A. M. Software Fault Interactions and Implications for Software Testing. IEEE TSE 30(6), 2004.
  7. Lehman, J., Stanley, K. O. Abandoning Objectives: Evolution Through the Search for Novelty Alone. Evolutionary Computation 19(2), 2011.
  8. Zheng, L. et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023.
  9. Gao, L., Schulman, J., Hilton, J. Scaling Laws for Reward Model Overoptimization. ICML 2023.
  10. Schulman, J. et al. Proximal Policy Optimization Algorithms. 2017.
  11. Noukhovitch, M. et al. Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models. ICLR 2025.

What to run next

recipes/01-simulate/bring-your-own-agent runs these eight steps on a callable of your own, offline and in seconds.
Last modified on September 22, 2026