fr-bench-pdf2md - Handwritten: leaderboard

Metric: Unit-test pass rate (%) of the model's Markdown conversion on fr-bench-pdf2md's handwritten (historical manuscripts and contemporary handwritten notes) pages (difficult French PDF pages from CCPDF and Gallica, chosen where two OCR models disagreed most), checking text presence, reading order and table structure after category-specific normalization; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 17 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Pro (Preview)60#64
2Gemini 3 Flash (Preview)58.2#78
3GPT-5.220.6#105
4GPT-5 Mini18.2#176
5Gemini 2.5 Flash Lite12.7#413

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=fr-bench-pdf2md-handwritten · How It Works · Data refreshed daily, snapshot 2026-10-11.