PureDocBench - Digital Degraded: leaderboard

Metric: Overall parsing score (0-100): mean of text accuracy (100 times one minus the text normalized edit distance), formula CDM and table TEDS, on the digitally degraded track (the same pages under ten simulated scene templates such as aging, binding, photocopying, compression, lighting and geometric artifacts) of 1,475 source pages in 10 domains and 66 subcategories, rendered from LLM-written HTML/CSS with source-linked annotations for text, formulas, tables and reading order; higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 58 models tracked.

Top models

#ModelScore
1GLM-5.3 Flash81.03
2Gemini 3.6 Flash79.11
3Qwen 3.8 27B78.86
4Qwen 3.5 122B A10B76.34
5Qwen 3.5 9B73.34
6Claude Opus 572.94
7Qwen 3.5 4B72.53
8GPT-5.6 Sol71.73
9Qwen3.8-Flash-Next70.96
10Qwen 3.5 27B70.73
11Kimi K2.6 (Non-reasoning)69.95
12Gemini 3.1 Pro (Preview)69.28
13Qwen 3.6 35B A3B69.16
14Qwen 3.5 397B A17B68.34
15Qwen 3.5 35B A3B68.04

Interactive version: theaggregate.ai/benchmark?slug=puredocbench-digital-degraded · How It Works · Data refreshed daily, snapshot 2026-10-07.