fr-bench-pdf2md - Long Tables: leaderboard

Metric: Unit-test pass rate (%) of the model's Markdown conversion on fr-bench-pdf2md's long-table (dense tables spanning the page) pages (difficult French PDF pages from CCPDF and Gallica, chosen where two OCR models disagreed most), checking text presence, reading order and table structure after category-specific normalization; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 17 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Flash (Preview)82.8#78
2Gemini 3 Pro (Preview)81.3#64
3GPT-5.273.2#105
4GPT-5 Mini66#176
5Gemini 2.5 Flash Lite20.5#413

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=fr-bench-pdf2md-long-tables · How It Works · Data refreshed daily, snapshot 2026-10-11.