DocScope - Region Grounding: leaderboard

Metric: Strict region grounding F1 (%): a GPT-5.5 multimodal judge labels each gold evidence region on the correctly retrieved pages as covered, imprecise or not covered by the predicted boxes, counting only covered over the answerable questions of the 730-question DocScope test split (human-annotated questions on long, visually rich PDF documents; the model receives the whole document and returns evidence pages, evidence regions, supporting facts and an answer), each stage scored only on the gold pages the model retrieved; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 18 models tracked.

Top models

#ModelScore
1GPT-5.457.4
2Claude Sonnet 4.644.5
3Claude Opus 4.742.9
4Gemini 3.1 Pro (Preview)39.7
5Gemini 3.1 Flash Lite36.2
6Qwen 3.6 Plus29.6
7Qwen 3.5 27B28.8
8Gemma 4 31B27.3
9Qwen 3 VL 235B A22B27.2
10Qwen 3.5 397B A17B24.3
11Gemma 4 26B A4B7.5

Interactive version: theaggregate.ai/benchmark?slug=docscope-region-grounding · How It Works · Data refreshed daily, snapshot 2026-10-07.