DocScope - Region Grounding: leaderboard
Metric: Strict region grounding F1 (%): a GPT-5.5 multimodal judge labels each gold evidence region on the correctly retrieved pages as covered, imprecise or not covered by the predicted boxes, counting only covered over the answerable questions of the 730-question DocScope test split (human-annotated questions on long, visually rich PDF documents; the model receives the whole document and returns evidence pages, evidence regions, supporting facts and an answer), each stage scored only on the gold pages the model retrieved; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 18 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 | 57.4 |
| 2 | Claude Sonnet 4.6 | 44.5 |
| 3 | Claude Opus 4.7 | 42.9 |
| 4 | Gemini 3.1 Pro (Preview) | 39.7 |
| 5 | Gemini 3.1 Flash Lite | 36.2 |
| 6 | Qwen 3.6 Plus | 29.6 |
| 7 | Qwen 3.5 27B | 28.8 |
| 8 | Gemma 4 31B | 27.3 |
| 9 | Qwen 3 VL 235B A22B | 27.2 |
| 10 | Qwen 3.5 397B A17B | 24.3 |
| 11 | Gemma 4 26B A4B | 7.5 |
Interactive version: theaggregate.ai/benchmark?slug=docscope-region-grounding · How It Works · Data refreshed daily, snapshot 2026-10-07.