The scripts are in the repository, not in the installed package. Clone it,
then
cd recipes/papers/context-lm before running the commands below. Browse this recipe on GitHub.0.25 x clip((mean success cost - c_i) / mean success cost, -1, 1) to the advantage of every context edit of a successful trajectory. Harness, task, reward, model, steps and seeds are the same in both arms.
Recipe
- Task: a seeded key-value log (ContextBench’s KV Store, cut down [1]). 5 chunks of 8
set key = valuelines over 8 keys; the asked key is set at least twice. Each step the model seescontext.mdand one chunk, then the chunk is gone, and it replies with the whole newcontext.md. The last step shows onlycontext.mdand the question; the model boxes a number. 512 train logs, 200 held out from a disjoint seed range. - Base:
Qwen/Qwen2.5-1.5B-Instruct. Costc_i: prefix-reuse tokens, each step’s prompt past the prefix it shares with the previous step’s prompt and reply, plus the reply (the token count the paper’skv_cache_flops.pyturns into FLOPs). - Baseline, stepwise GRPO: every step of a trajectory gets reward minus the mean over its group of 8.
- Recipe: baseline plus
w_eff = 0.25times Eq. 6 on the five context edits, not on the answer step (the paper’s edit mask). A group with fewer than 2 successes gets no efficiency credit. - Both arms: TRL GRPO machinery + LoRA r=32, 60 steps, 4 logs x 8 trajectories a step, lr 1e-4, on-policy, no KL, seeds 17 and 18.
- Eval: 4 trajectories on each of the 200 held-out logs. pass@1 with a paired 95% interval (
wai.compare), and tokens per trajectory, paired by log.
Run
Result
Recipe vs baseline on pass@1, seed 17 paired (
wai.compare, train_runs= both seeds): -0.22 [-0.29, -0.16]. Verdict: flat. The two recipe seeds land on opposite sides of the baseline, so the training-seed spread swallows the gap.
Tokens a trajectory, recipe minus baseline, paired by held-out log: seed 17 +181 [+172, +189] (+17%), seed 18 -318 [-330, -307] (-23%).
The two seeds learned different files. On seed 18 the recipe did what the paper says: a bare pearl: 231 table, one line a key, the cheapest file of any arm. Accuracy went up to 0.96 and tokens went down 23%. On seed 17 it locked onto copying the latest chunk verbatim, which answers only when the asked key’s last set falls in that chunk. That happens on about 70% of logs, and the arm scored 0.70. The baseline seeds also differ, but both still answer: seed 17 folds the log into one line of current values, and seed 18 copies the last two chunks forward at 308 characters.
Checks
Every number in this section is read fromresults.json.
Climb
Learned
- Eq. 6 can reach the paper’s result at 1.5B. On seed 18 it gave a smaller file, 23% fewer tokens and higher accuracy than the baseline. It did so on one seed of two, though.
- An efficiency term that only pays successes still rewards a cheap partial success. Copying the last chunk wins about 70% of logs, and a seed that finds it early can keep it. The paper’s run had a turn budget, a
shrinkedit gate and an LLM judge; none of those is here to push back. - Next: three or more seeds per arm, and a check on the files rather than the answers, such as the share of current values
context.mdstill holds after each chunk. That would see the shortcut before the holdout does.w_eff0.1 is the cheap ablation.
References
- Context Language Models. arXiv:2609.37725, 2026. Code: https://github.com/facebookresearch/context-language-models
- Shao, Z. et al. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv:2402.03300, 2024.
- Lambert, N. Reinforcement Learning from Human Feedback. arXiv:2504.12501, 2025. Chapter Evaluation.
- Gao, L., Schulman, J., Hilton, J. Scaling Laws for Reward Model Overoptimization. ICML 2023. arXiv:2210.10760.