The scripts are in the repository, not in the installed package. Clone it,
then
cd recipes/04-train/parsebench before running the commands below. Browse this recipe on GitHub.python selftest.py needs nothing.
Results
Overall is the mean of the five ParseBench dimensions, the number the leaderboard reports. Intervals are 95% percentile bootstraps over documents, 2,000 draws, frommodal run serve/bench.py::bands. Every number here is in
results.json.
Leaderboard rows are from ParseBench’s
leaderboard.csv, read 2026-09-25. The
base model on its own reproduces its leaderboard row on our dev split (70.84
against 70.79), so the gain is the harness, not a different server.
The honesty rules
ParseBench ships no training split, so every number above could be a memorized benchmark if we were careless. Three rules keep it honest:- No ParseBench page is ever trained on. The RL data is ChartNet [3] (CDLA-Permissive-2.0), and any ChartNet chart whose values cover 30% or more of a ParseBench chart page’s rules is dropped (281 of 29.7k).
- Tune on dev only.
waiparse/split.pyhashes the source report name (the page suffix_pNremoved) and sends one report in five to dev. Every page of a report lands on the same side, so a report’s fonts and layout cannot leak from dev into test. - Report test and full. Test is the other four reports in five, never
used to pick a setting. Full is for leaderboard parity. One piece of the
headline, the detector chart trigger in v4, was motivated by a failure seen
on test (a chart the layout pass read as bare label lines). Its settings
were tuned on dev only, and
waiparse/pipelines.pysays so.
Run it
Needs: Python 3.11+,pip install modal, modal setup, and these in your
own Modal workspace:
DOCPARSE_WORKSPACE
(waiparse/endpoints.py). DOCPARSE_SERVER, DOCPARSE_LAYOUT_URL and
DOCPARSE_TUNED_SERVER override one URL each. Modal shortens a label longer
than 63 characters, so with a long workspace name, copy the URL modal deploy
printed.
From this directory:
/data/split_dev or /data/split_test from
the full set. Results land in the docparse-runs volume at
/runs/bench/<pipeline>/<split>. --group table (or chart, text,
layout) runs one dimension. Name a function after bench.py::: a bare
modal run serve/bench.py is ambiguous.
What each harness piece bought
The agent (waiparse/agent.py) runs per page: a layout pass, then specialist
passes on regions, then emits markdown and layout boxes in the shape the
scorers read. Each row is one change, scored on dev against the row above it.
On test, the same pieces hold:
wai_agent_full 77.12, and wai_agent_v4
(full, plus three-sample voting on tables as well as charts, plus the detector’s
chart label as a second trigger for the page chart pass) 78.53, a paired +1.41
(0.87 to 2.01). wai_agent_v5 (v4 with medium chart effort) won on dev
charts, 86.7 against 83.5 to 85.2 at xhigh, but scored 78.19 on test, a paired
-0.34 (-0.80 to 0.10): no difference, so v4 stays the headline.
Dev runs are one pass each. Tables swung 76 to 81 between single-sample runs,
which is why voting is on for tables in v4. The gap from dev to test (79.97 to
77.12 for wai_agent_full) is mostly layout and charts; the detector
thresholds were tuned on dev.
Detector ablations run without new model calls: serve/relayout.py attaches
detector boxes to existing raw outputs and re-scores the layout group
(wai_agent_det_* in waiparse/pipelines.py).
What did not work
- Chart RL on crops.
train/configs/charts-v1.toml: LoRA r32 on Qwen3.8-27B, GRPO advantages [5] with the CISPO loss [6] as ScaleRL runs it [7], reward = the share of ParseBench chart rules the reply passes, in the spirit of olmOCR 2’s unit-test rewards [8]. Validation reward rose from 0.569 to 0.624 by step 100, but ParseBench dev charts fell to 82.1 against the base’s 83.5 to 85.2. - Chart RL on whole pages.
train/configs/charts-pages-v1.tomltrains on synthetic report pages composed from ChartNet charts with the page chart prompt the agent uses. Validation reward rose from 0.539 to 0.559 by step 50; ParseBench dev charts at medium effort were 85.9 against the base’s 86.7. - Medium chart effort won on dev and did not hold on test (above).
- A lower detector threshold keeps raising the layout score (0.1 gives 83.5 dev layout at 52 boxes a page against 35 in the ground truth) because the headline has no false-positive penalty. We stayed at 0.2 on purpose.
train/README.md has the launcher, the
checks to run before a launch, and how to resume and serve an adapter.
Where to hill-climb next
- Tables. 80.5 here against 94.3 for Claude Opus 5.5. The largest gap to the best closed model.
- Formatting. 70.3, the lowest dimension.
modal run serve/bench.py::fmt_failures --spec wai_agent_v4:devprints the failing rules with the markdown around each. - The dev to test drop in layout and charts. Tune the detector knobs
(
det_*inagent.DEFAULTS) with an interval, not a single dev run. - Chart data that looks like report charts for RL: rendered charts with styled legends, data labels and small multiples, or rewards on real unlabeled pages through self-consistency.
Cost and time
Modal on-demand prices (modal.com/pricing, read 2026-09-28): H200 0.80 an hour. The vLLM app runs up to four H200 containers (DOCPARSE_MAX_CONTAINERS) and scales to zero after ten idle minutes; the
detector runs up to four L4s; the benchmark runner is a CPU container. A
single-dimension dev run is the cheap loop; test and full are the expensive
ones. RL runs on four H200s (50;
charts-pages-v1 took about 3 minutes a step, about $45 for 50 steps.
Traps
- Reasoning parser. vLLM needs
--reasoning-parser qwen3, or the thinking text lands in the content and breaks the JSON the layout pass returns. - Model name. ParseBench’s Qwen layout adapter keys on the served model
name starting with
qwen3., hence--served-model-name qwen3.8-27b. - pdfium is not thread-safe. ParseBench runs documents in threads.
waiparse/render.pytakes a global lock and closes bitmaps and pages inside it; left to the garbage collector, the full-set run segfaulted. - Image size. The chart RL stack starved at 2048 px pages (over half the rollouts timed out); page tasks are 1400 px. The agent reads charts at 2048 px (86.7 dev charts against 82.5 at 1400), so evaluate an adapter at 2048.
- Weight broadcast. Loading thousands of inlined page images outlasts
prime-rl’s default 1200 s broadcast wait. Put
[weight_broadcast] timeout = 3600at the top level of the config: prime-rl rebuilds the trainer and orchestrator sections from it and overwrites their own timeouts. Capenv.taskset.max_taskstoo. - Router queue. At 1,024 in-flight rollouts the vllm-router queue (100 requests, 60 s) expired about half of them as “Request timed out”; the page config uses 128.
- Long requests. A Modal web endpoint answers a long request with a 303;
follow redirects (
curl -L). modal volume getwith an empty path downloads the whole volume.- Use
--spawnfor RL, notmodal run --detach. A detached run dies with the local client.
References
- Zhang et al. ParseBench: A Document Parsing Benchmark for AI Agents. arXiv:2604.08538, 2026.
- Cui et al. PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing. arXiv:2601.21957, 2026. PP-DocLayoutV3 is its layout stage; weights at
PaddlePaddle/PP-DocLayoutV3_safetensors, Apache-2.0. - Kondic et al. ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding. arXiv:2603.27064, 2026.
- Wang et al. Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv:2203.11171, 2022.
- Shao et al. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv:2402.03300, 2024.
- MiniMax. MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention. arXiv:2506.13585, 2025.
- Khatri et al. The Art of Scaling Reinforcement Learning Compute for LLMs. arXiv:2510.13786, 2025.
- Poznanski et al. olmOCR 2: Unit Test Rewards for Document OCR. arXiv:2510.19817, 2025.