Skip to main content
The scripts are in the repository, not in the installed package. Clone it, then cd recipes/papers/ptgs before running the commands below. Browse this recipe on GitHub.
The answer: the temperature fix tied plain RL. Practicing more won, at almost 4x the compute. No method lost range. Three ways to practice: plain RL draws 4 attempts at temperature 0.9, the temperature fix draws 4 attempts hotter on hard problems and cooler on easy ones, practice-more keeps drawing up to 32 and trains on 4

Result

First-try and 8-try accuracy on 320 held-out MATH problems: plain RL 29% and 56%, temperature fix 31% and 60%, practice more 35% and 64%, untrained 20% and 50%
  • Temperature fix vs plain RL: +0.4 points first try, range -1.7 to +2.5. Not proven.
  • Practice more vs plain RL: +3.1 points, range +0.7 to +5.4. Both its runs beat both plain runs.
Practice that taught nothing: plain RL 68%, temperature fix 71%, practice more 41%. Minutes per run: 52, 52, 196
  • Why the fix did not help: heating a problem the model fails rarely makes it pass, and cooling an easy one makes every attempt pass. Both teach nothing.
Sharpening Tax at 8 tries with 95% ranges: every arm crosses zero
  • No tax here: 80 small training steps on a 1.5B model did not sharpen it enough. The paper sees the tax on fully post-trained models at up to 128 tries.

The paper

Paper: Sharpening Tax in Post-Training, Changdae Oh et al., arXiv:2610.01509, October 2026. https://arxiv.org/abs/2610.01509 Book: RL learns by comparing attempts at one problem, so a problem whose attempts all pass or all fail teaches nothing [1][2]. Claim: sampling hard problems hotter and easy ones cooler during training raises first-try accuracy without shrinking what the model can solve in many tries [7]. The change: each problem’s 4 training attempts are drawn at their own temperature (0.6 to 1.35) instead of 0.9 for all. A third arm, Reinforce-Ada, draws more attempts instead [3].

Recipe

Run

Checks

Climb

Learned

  • Hard problems needed more attempts, not hotter ones.
  • The fix costs nothing extra and plugs into TRL in about 40 lines.
  • Next: the paper’s GRPO setting (tau 1.4) and 64 attempts per problem, so the tax has room to show.
Verified 2026-10-02, whileai 0.127, TRL 0.19.1. 804 GPU minutes, $26.79 on L40S. Run page: https://while.ai/platform/training/run_ef3d2f55c789144d

References

  1. Shao, Z. et al. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv:2402.03300, 2024.
  2. Lambert, N. Reinforcement Learning from Human Feedback. arXiv:2504.12501, 2025. Chapter Reinforcement Learning.
  3. Xiong, W. et al. Reinforce-Ada: An Adaptive Sampling Framework under Non-linear RL Objectives. arXiv:2510.04996, 2025.
  4. Lambert, N. Reinforcement Learning from Human Feedback. arXiv:2504.12501, 2025. Chapter Evaluation.
  5. Lambert, N. et al. Tülu 3: Pushing Frontiers in Open Language Model Post-Training. arXiv:2411.15124, 2024.
  6. Gao, L., Schulman, J., Hilton, J. Scaling Laws for Reward Model Overoptimization. ICML 2023. arXiv:2210.10760.
  7. Oh, C. et al. Sharpening Tax in Post-Training. arXiv:2610.01509, 2026.
Last modified on October 2, 2026