PureDocBench: leaderboard

Metric: Three-track average (0-100): mean of the clean, digitally degraded and real degraded track Overall scores, each the mean of text accuracy (100 times one minus the text normalized edit distance), formula CDM and table TEDS, on 1,475 source pages in 10 domains and 66 subcategories, rendered from LLM-written HTML/CSS with source-linked annotations for text, formulas, tables and reading order; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 58 models tracked.

Top models

#ModelScore
1GLM-5.3 Flash79.69
2Gemini 3.6 Flash78.27
3Qwen 3.8 27B77.63
4Claude Opus 577.07
5Qwen 3.5 122B A10B74.11
6Qwen 3.5 9B70.89
7Gemini 3.1 Pro (Preview)70.43
8Kimi K2.6 (Non-reasoning)70.1
9Qwen 3.5 4B69.82
10Qwen 3.5 27B69.57
11GPT-5.6 Sol69.16
12Qwen3.8-Flash-Next68.54
13Qwen 3.6 35B A3B67.14
14Qwen 3.5 397B A17B66.72
15Qwen 3.5 35B A3B65.68

Interactive version: theaggregate.ai/benchmark?slug=puredocbench · How It Works · Data refreshed daily, snapshot 2026-10-07.