Dr. DocBench - Reading Order Edit Distance: leaderboard

Metric: Normalized edit distance of the predicted block reading order against the expert order, times 100 (0-100), 4,514 expert-annotated difficult pages sampled by parser failure from long books in 52 subject domains; each page parsed with the model's own document-to-markdown prompt (specialized parsers through their official APIs); lower is better. Source: arxiv.org. Saturation forecast: Around December 2026. 12 models tracked.

Top models

#ModelScore
1GPT-5.517
2Kimi K2.518
3Claude Opus 4.619
4Gemini 3.1 Pro (Preview)21
5Qwen 3.5 122B A10B23
6Qwen 3.5 Plus26
7Qwen 3.5 Flash26
8GPT-4o31
9Nemotron Nano 12B v2 VL56.4

Interactive version: theaggregate.ai/benchmark?slug=dr-docbench-reading-order-edit-distance · How It Works · Data refreshed daily, snapshot 2026-09-29.