VisReason (Vision-Centric) - Localized Reasoning: leaderboard

Metric: Accuracy (%) on the Localized Reasoning category (perceptual reasoning: infer the target instance in an image and localize it with a bounding box) of VisReason, answers scored by format (regular-expression match, GPT-5-mini judge for open-ended answers, IoU above 0.5 for boxes), zero-shot with chain-of-thought prompt templates, one per answer format; rows not shaded gray in the paper use explicit reasoning (GPT-5 and Gemini 3 Pro at low reasoning effort); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 22 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview) (Low)44.2
2GPT-5.2 (Thinking)24.6
3GPT-5 (Low)15
4GPT-5 Mini (High)12.5
5GPT-5 Mini (Low)12.5
6Qwen 3 VL 8B (Thinking)9.6
7GPT-5 Mini (Medium)7.5
8Qwen 3 VL 235B A22B (Thinking)7.1
9GPT-4o6.3
10GPT-5 Nano3.8
11O4 Mini3.8
12Qwen 2.5 VL 32B Instruct3.3
13InternVL3-14B1.3
14Qwen 3 VL 8B Instruct1.3
15Qwen 3 VL 32B (Thinking)1.3

Interactive version: theaggregate.ai/benchmark?slug=visreason-vision-centric-localized-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-07.