PureDocBench - Digital Degraded - Formula CDM: leaderboard
Metric: Formula Character Detection Matching score (0-100) between the rendered predicted and reference LaTeX formulas, on the digitally degraded track (the same pages under ten simulated scene templates such as aging, binding, photocopying, compression, lighting and geometric artifacts) of 1,475 source pages in 10 domains and 66 subcategories, rendered from LLM-written HTML/CSS with source-linked annotations for text, formulas, tables and reading order; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 58 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.6 Flash | 78.88 |
| 2 | GLM-5.3 Flash | 78.51 |
| 3 | Claude Opus 5 | 73.43 |
| 4 | Qwen3.8-Flash-Next | 72.87 |
| 5 | GPT-5.6 Sol | 72.83 |
| 6 | Qwen 3.8 27B | 71.98 |
| 7 | Qwen 3.5 4B | 68.88 |
| 8 | Qwen 3.5 122B A10B | 67.82 |
| 9 | Qwen 3.6 35B A3B | 67.67 |
| 10 | Qwen 3.5 9B | 67 |
| 11 | Qwen 3.6 27B | 66.54 |
| 12 | Gemini 3.1 Pro (Preview) | 65.81 |
| 13 | Qwen 3.5 35B A3B | 64.78 |
| 14 | Kimi K2.6 (Non-reasoning) | 64.69 |
| 15 | Qwen 3.5 27B | 64.61 |
Interactive version: theaggregate.ai/benchmark?slug=puredocbench-digital-degraded-formula-cdm · How It Works · Data refreshed daily, snapshot 2026-10-07.