LED (Image + JSON) - Element Error F1: leaderboard
Metric: Element-level macro F1 (x 100): labelling each predicted box with one of eight structural error types or none, given the page image plus the predicted layout as JSON (with ground-truth boxes), on LED's 4,996 DocLayNet test pages with structural layout errors (missing, hallucinated, resized, misclassified, split, merged, overlapping, duplicated regions) injected at the rates a commercial layout model produces, zero-shot via OpenRouter at temperature 1.0; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Gemini 2.5 Pro | 44.3 | #145 |
| 2 | Gemini 2.5 Flash | 33.3 | #237 |
| 3 | GPT-4o Mini | 15.9 | #588 |
| 4 | Gemini 2.5 Flash Lite | 12.7 | #413 |
| 5 | Llama 4 Maverick | 7.5 | #451 |
| 6 | GPT-4o | 6.6 | #333 |
| 7 | Llama 4 Scout | 0.2 | #646 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=led-image-plus-json-element-error-f1 · How It Works · Data refreshed daily, snapshot 2026-10-11.