The scripts are in the repository, not in the installed package. Clone it,
then
cd recipes/papers/ptgs before running the commands below. Browse this recipe on GitHub.Result
- Temperature fix vs plain RL: +0.4 points first try, range -1.7 to +2.5. Not proven.
- Practice more vs plain RL: +3.1 points, range +0.7 to +5.4. Both its runs beat both plain runs.
- Why the fix did not help: heating a problem the model fails rarely makes it pass, and cooling an easy one makes every attempt pass. Both teach nothing.
- No tax here: 80 small training steps on a 1.5B model did not sharpen it enough. The paper sees the tax on fully post-trained models at up to 128 tries.
The paper
Paper: Sharpening Tax in Post-Training, Changdae Oh et al., arXiv:2610.01509, October 2026. https://arxiv.org/abs/2610.01509 Book: RL learns by comparing attempts at one problem, so a problem whose attempts all pass or all fail teaches nothing [1][2]. Claim: sampling hard problems hotter and easy ones cooler during training raises first-try accuracy without shrinking what the model can solve in many tries [7]. The change: each problem’s 4 training attempts are drawn at their own temperature (0.6 to 1.35) instead of 0.9 for all. A third arm, Reinforce-Ada, draws more attempts instead [3].Recipe
Run
Checks
Climb
Learned
- Hard problems needed more attempts, not hotter ones.
- The fix costs nothing extra and plugs into TRL in about 40 lines.
- Next: the paper’s GRPO setting (tau 1.4) and 64 attempts per problem, so the tax has room to show.
References
- Shao, Z. et al. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv:2402.03300, 2024.
- Lambert, N. Reinforcement Learning from Human Feedback. arXiv:2504.12501, 2025. Chapter Reinforcement Learning.
- Xiong, W. et al. Reinforce-Ada: An Adaptive Sampling Framework under Non-linear RL Objectives. arXiv:2510.04996, 2025.
- Lambert, N. Reinforcement Learning from Human Feedback. arXiv:2504.12501, 2025. Chapter Evaluation.
- Lambert, N. et al. Tülu 3: Pushing Frontiers in Open Language Model Post-Training. arXiv:2411.15124, 2024.
- Gao, L., Schulman, J., Hilton, J. Scaling Laws for Reward Model Overoptimization. ICML 2023. arXiv:2210.10760.
- Oh, C. et al. Sharpening Tax in Post-Training. arXiv:2610.01509, 2026.