DocAttriBench - Answer Locating: leaderboard

Metric: Box F1 (%; answer locating: given only the question, predict the evidence boxes, mean over the seven test sets DocVQA, VisualMRC, VISA, SlideVQA, VisualWebBench, LongDocURL and MMLongBench-Doc; zero-shot, with model-specific prompts and box parsers; element-level evidence boxes from the MAPPET attribution pipeline over Docling layout regions, a box counting when IoU >= 0.5). Source: arxiv.org. Saturation forecast: Around December 2026. 16 models tracked.

Top models

#ModelScore
1Qwen 3 VL 8B21.1
2Qwen 2.5 VL 32B Instruct19.3
3InternVL3-38B18.6
4Qwen 2.5 VL 7B Instruct17.9
5InternVL3-8B11.1
6InternVL2.5-2B0.4

Interactive version: theaggregate.ai/benchmark?slug=docattribench-answer-locating · How It Works · Data refreshed daily, snapshot 2026-09-26.