PureDocBench - Clean - Formula CDM: leaderboard

Metric: Formula Character Detection Matching score (0-100) between the rendered predicted and reference LaTeX formulas, on the clean track (pages rendered straight from their HTML/CSS source) of 1,475 source pages in 10 domains and 66 subcategories, rendered from LLM-written HTML/CSS with source-linked annotations for text, formulas, tables and reading order; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 58 models tracked.

Top models

#ModelScore
1GLM-5.3 Flash80.16
2Gemini 3.6 Flash78.29
3Claude Opus 575.92
4GPT-5.6 Sol74.96
5Qwen3.8-Flash-Next72.28
6Qwen 3.8 27B71.68
7Qwen 3.6 35B A3B70.39
8Qwen 3.5 4B69.96
9Qwen 3.6 27B69.36
10Qwen 3.5 122B A10B67.96
11Qwen 3.5 9B67.6
12Kimi K2.6 (Non-reasoning)66.93
13Qwen 3.5 27B66.36
14Gemini 3.1 Pro (Preview)65.63
15Qwen 3.5 397B A17B65.26

Interactive version: theaggregate.ai/benchmark?slug=puredocbench-clean-formula-cdm · How It Works · Data refreshed daily, snapshot 2026-10-07.