wai.methods.ContextFile is that harness, plus the credit the paper trains it with.
The episode
An episode is edits and one answer. At edit the model sees the file and the next input, then writes the whole new file; the input is gone after that turn. At the end it sees only the file and the question.wai.methods.KVLog is a seeded task for it: 5 chunks of 8 set key = value lines over 8 keys, and a question about one key’s final value.
The credit
Stepwise GRPO gives every step of trajectory the outcome advantage over its group. The paper adds a success-gated efficiency advantage on the edits (Eq. 6): among the group’s successes, a trajectory cheaper than their mean goes up and a dearer one goes down, clipped to . Failures get nothing, and so does every trajectory in a group with fewer than two successes. Cost is prefix-reuse tokens: each step pays for the prompt past what is still in the KV cache, plus its reply.Train it
clm.trainer(GRPOTrainer) is a TRL GRPOTrainer whose generation step plays the episode and gives each step its credit. Pass env= when you construct it, and set num_iterations=1 and beta=0.
clm.play(generate, tasks, env) runs episodes with any generator that maps a list of chats and a token cap to Reply objects. clm.report(episodes) prints the pass rate, tokens a trajectory and file size.
What it did
recipes/papers/context-lm ran Qwen2.5-1.5B-Instruct for 60 steps with LoRA. It evaluated on 200 held-out logs, four seeds an arm, three arms:
Every seed learned to keep its context. Eq. 6 added +0.02 [+0.01, +0.03] pass@1 at 15% fewer tokens, and the verdict is flat because the seeds of plain GRPO spread wider than that. The trained files are short
key: value tables, not copies of the log.
References
- Context Language Models. arXiv:2609.37725, 2026. https://github.com/facebookresearch/context-language-models
- Shao, Z. et al. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv:2402.03300, 2024.