fr-bench-pdf2md - Graphics: leaderboard

Metric: Unit-test pass rate (%) of the model's Markdown conversion on fr-bench-pdf2md's graphics (scientific figures and charts whose embedded text must be transcribed) pages (difficult French PDF pages from CCPDF and Gallica, chosen where two OCR models disagreed most), checking text presence, reading order and table structure after category-specific normalization; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 17 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5 Mini81.6#176
2GPT-5.280.2#105
3Gemini 3 Pro (Preview)77.3#64
4Gemini 3 Flash (Preview)73.4#78
5Gemini 2.5 Flash Lite42.2#413

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=fr-bench-pdf2md-graphics · How It Works · Data refreshed daily, snapshot 2026-10-11.