PureDocBench - Real Degraded: leaderboard

Metric: Overall parsing score (0-100): mean of text accuracy (100 times one minus the text normalized edit distance), formula CDM and table TEDS, on the real degraded track (the same pages through four physical acquisition chains: phone photos of prints and of photocopies, screen photographs, and screenshots with social-media compression) of 1,475 source pages in 10 domains and 66 subcategories, rendered from LLM-written HTML/CSS with source-linked annotations for text, formulas, tables and reading order; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 58 models tracked.

Top models

#ModelScore
1Gemini 3.6 Flash75.73
2Claude Opus 575.27
3GLM-5.3 Flash74.99
4Qwen 3.8 27B73.59
5Gemini 3.1 Pro (Preview)71.98
6Qwen 3.5 122B A10B69.85
7Kimi K2.6 (Non-reasoning)68.02
8Qwen 3.5 27B65.92
9Qwen 3.5 9B65.45
10Qwen 3.5 4B63.47
11Qwen 3.5 397B A17B62.7
12Qwen3.8-Flash-Next62.42
13GPT-5.6 Sol61.59
14Qwen 3.5 35B A3B60.59
15Qwen 3.6 35B A3B60.12

Interactive version: theaggregate.ai/benchmark?slug=puredocbench-real-degraded · How It Works · Data refreshed daily, snapshot 2026-10-07.