LED (Image + JSON) - Error Detection: leaderboard
Metric: Document-level error detection accuracy (%): whether the predicted layout of a page contains any structural error, given the page image plus the predicted layout as JSON, on LED's 4,996 DocLayNet test pages with structural layout errors (missing, hallucinated, resized, misclassified, split, merged, overlapping, duplicated regions) injected at the rates a commercial layout model produces, zero-shot via OpenRouter at temperature 1.0; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 8 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Gemini 2.5 Pro | 63.6 | #145 |
| 2 | Gemini 2.5 Flash | 61 | #237 |
| 3 | GPT-4o | 59.7 | #333 |
| 4 | GPT-4o Mini | 53.8 | #588 |
| 5 | Llama 4 Maverick | 47.6 | #451 |
| 6 | Llama 4 Scout | 46.1 | #646 |
| 7 | Gemini 2.5 Flash Lite | 42.1 | #413 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=led-image-plus-json-error-detection · How It Works · Data refreshed daily, snapshot 2026-10-11.