Skip to main content
The scripts are in the repository, not in the installed package. Clone it, then cd recipes/04-train/parsebench before running the commands below. Browse this recipe on GitHub.
An open-weight agent that turns PDF pages into markdown, HTML tables and layout boxes, scored on ParseBench [1]. It is Qwen3.8-27B served by vLLM plus the PP-DocLayoutV3 layout detector [2], both on Modal in your own workspace. On ParseBench it scores 78.89 on the full set and 78.53 on held-out test documents. That puts while.ai first among open-weight entries on the leaderboard (rakedoc-nano, 77.23, sits below the full-set interval), and Claude Opus 5.5 at high effort (79.85) sits inside it. What you will learn: how to hill-climb an agent harness on a public benchmark without fooling yourself (a document-level dev/test split, tune on dev, report test and full), which harness pieces moved the score and by how much, and why two RL runs on synthetic chart data did not transfer. You need a Modal account with an H200 quota, a Hugging Face token, and a few GPU hours. python selftest.py needs nothing.

Results

Overall is the mean of the five ParseBench dimensions, the number the leaderboard reports. Intervals are 95% percentile bootstraps over documents, 2,000 draws, from modal run serve/bench.py::bands. Every number here is in results.json. Leaderboard rows are from ParseBench’s leaderboard.csv, read 2026-09-25. The base model on its own reproduces its leaderboard row on our dev split (70.84 against 70.79), so the gain is the harness, not a different server.

The honesty rules

ParseBench ships no training split, so every number above could be a memorized benchmark if we were careless. Three rules keep it honest:
  1. No ParseBench page is ever trained on. The RL data is ChartNet [3] (CDLA-Permissive-2.0), and any ChartNet chart whose values cover 30% or more of a ParseBench chart page’s rules is dropped (281 of 29.7k).
  2. Tune on dev only. waiparse/split.py hashes the source report name (the page suffix _pN removed) and sends one report in five to dev. Every page of a report lands on the same side, so a report’s fonts and layout cannot leak from dev into test.
  3. Report test and full. Test is the other four reports in five, never used to pick a setting. Full is for leaderboard parity. One piece of the headline, the detector chart trigger in v4, was motivated by a failure seen on test (a chart the layout pass read as bare label lines). Its settings were tuned on dev only, and waiparse/pipelines.py says so.

Run it

Needs: Python 3.11+, pip install modal, modal setup, and these in your own Modal workspace:
Every server URL the agent calls is built from DOCPARSE_WORKSPACE (waiparse/endpoints.py). DOCPARSE_SERVER, DOCPARSE_LAYOUT_URL and DOCPARSE_TUNED_SERVER override one URL each. Modal shortens a label longer than 63 characters, so with a long workspace name, copy the URL modal deploy printed. From this directory:
The first run on a split writes /data/split_dev or /data/split_test from the full set. Results land in the docparse-runs volume at /runs/bench/<pipeline>/<split>. --group table (or chart, text, layout) runs one dimension. Name a function after bench.py::: a bare modal run serve/bench.py is ambiguous.

What each harness piece bought

The agent (waiparse/agent.py) runs per page: a layout pass, then specialist passes on regions, then emits markdown and layout boxes in the shape the scorers read. Each row is one change, scored on dev against the row above it. On test, the same pieces hold: wai_agent_full 77.12, and wai_agent_v4 (full, plus three-sample voting on tables as well as charts, plus the detector’s chart label as a second trigger for the page chart pass) 78.53, a paired +1.41 (0.87 to 2.01). wai_agent_v5 (v4 with medium chart effort) won on dev charts, 86.7 against 83.5 to 85.2 at xhigh, but scored 78.19 on test, a paired -0.34 (-0.80 to 0.10): no difference, so v4 stays the headline. Dev runs are one pass each. Tables swung 76 to 81 between single-sample runs, which is why voting is on for tables in v4. The gap from dev to test (79.97 to 77.12 for wai_agent_full) is mostly layout and charts; the detector thresholds were tuned on dev. Detector ablations run without new model calls: serve/relayout.py attaches detector boxes to existing raw outputs and re-scores the layout group (wai_agent_det_* in waiparse/pipelines.py).

What did not work

  • Chart RL on crops. train/configs/charts-v1.toml: LoRA r32 on Qwen3.8-27B, GRPO advantages [5] with the CISPO loss [6] as ScaleRL runs it [7], reward = the share of ParseBench chart rules the reply passes, in the spirit of olmOCR 2’s unit-test rewards [8]. Validation reward rose from 0.569 to 0.624 by step 100, but ParseBench dev charts fell to 82.1 against the base’s 83.5 to 85.2.
  • Chart RL on whole pages. train/configs/charts-pages-v1.toml trains on synthetic report pages composed from ChartNet charts with the page chart prompt the agent uses. Validation reward rose from 0.539 to 0.559 by step 50; ParseBench dev charts at medium effort were 85.9 against the base’s 86.7.
  • Medium chart effort won on dev and did not hold on test (above).
  • A lower detector threshold keeps raising the layout score (0.1 gives 83.5 dev layout at 52 boxes a page against 35 in the ground truth) because the headline has no false-positive penalty. We stayed at 0.2 on purpose.
Both RL runs say the same thing: synthetic ChartNet charts do not teach what ParseBench’s report charts need. train/README.md has the launcher, the checks to run before a launch, and how to resume and serve an adapter.

Where to hill-climb next

  • Tables. 80.5 here against 94.3 for Claude Opus 5.5. The largest gap to the best closed model.
  • Formatting. 70.3, the lowest dimension. modal run serve/bench.py::fmt_failures --spec wai_agent_v4:dev prints the failing rules with the markdown around each.
  • The dev to test drop in layout and charts. Tune the detector knobs (det_* in agent.DEFAULTS) with an interval, not a single dev run.
  • Chart data that looks like report charts for RL: rendered charts with styled legends, data labels and small multiples, or rewards on real unlabeled pages through self-consistency.

Cost and time

Modal on-demand prices (modal.com/pricing, read 2026-09-28): H200 4.54anhour,L44.54 an hour, L4 0.80 an hour. The vLLM app runs up to four H200 containers (DOCPARSE_MAX_CONTAINERS) and scales to zero after ten idle minutes; the detector runs up to four L4s; the benchmark runner is a CPU container. A single-dimension dev run is the cheap loop; test and full are the expensive ones. RL runs on four H200s (18.16anhour):‘charts−v1‘tookabout1.6minutesastep,soits100stepswereunderthreehoursandabout18.16 an hour): `charts-v1` took about 1.6 minutes a step, so its 100 steps were under three hours and about 50; charts-pages-v1 took about 3 minutes a step, about $45 for 50 steps.

Traps

  • Reasoning parser. vLLM needs --reasoning-parser qwen3, or the thinking text lands in the content and breaks the JSON the layout pass returns.
  • Model name. ParseBench’s Qwen layout adapter keys on the served model name starting with qwen3., hence --served-model-name qwen3.8-27b.
  • pdfium is not thread-safe. ParseBench runs documents in threads. waiparse/render.py takes a global lock and closes bitmaps and pages inside it; left to the garbage collector, the full-set run segfaulted.
  • Image size. The chart RL stack starved at 2048 px pages (over half the rollouts timed out); page tasks are 1400 px. The agent reads charts at 2048 px (86.7 dev charts against 82.5 at 1400), so evaluate an adapter at 2048.
  • Weight broadcast. Loading thousands of inlined page images outlasts prime-rl’s default 1200 s broadcast wait. Put [weight_broadcast] timeout = 3600 at the top level of the config: prime-rl rebuilds the trainer and orchestrator sections from it and overwrites their own timeouts. Cap env.taskset.max_tasks too.
  • Router queue. At 1,024 in-flight rollouts the vllm-router queue (100 requests, 60 s) expired about half of them as “Request timed out”; the page config uses 128.
  • Long requests. A Modal web endpoint answers a long request with a 303; follow redirects (curl -L).
  • modal volume get with an empty path downloads the whole volume.
  • Use --spawn for RL, not modal run --detach. A detached run dies with the local client.

References

  1. Zhang et al. ParseBench: A Document Parsing Benchmark for AI Agents. arXiv:2604.08538, 2026.
  2. Cui et al. PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing. arXiv:2601.21957, 2026. PP-DocLayoutV3 is its layout stage; weights at PaddlePaddle/PP-DocLayoutV3_safetensors, Apache-2.0.
  3. Kondic et al. ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding. arXiv:2603.27064, 2026.
  4. Wang et al. Self-Consistency Improves Chain of Thought Reasoning in Language Models. arXiv:2203.11171, 2022.
  5. Shao et al. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv:2402.03300, 2024.
  6. MiniMax. MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention. arXiv:2506.13585, 2025.
  7. Khatri et al. The Art of Scaling Reinforcement Learning Compute for LLMs. arXiv:2510.13786, 2025.
  8. Poznanski et al. olmOCR 2: Unit Test Rewards for Document OCR. arXiv:2510.19817, 2025.
Last modified on September 28, 2026