VisReason (Vision-Centric) - Board Reasoning: leaderboard

Metric: Accuracy (%) on the Board Reasoning category (structural reasoning: infer the correct outcome of a board-game position shown as an image) of VisReason, answers scored by format (regular-expression match, GPT-5-mini judge for open-ended answers, IoU above 0.5 for boxes), zero-shot with chain-of-thought prompt templates, one per answer format; rows not shaded gray in the paper use explicit reasoning (GPT-5 and Gemini 3 Pro at low reasoning effort); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 22 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview) (Low)51.1
2GPT-5 (Low)44.4
3GPT-5.2 (Thinking)43.7
4GPT-5 Mini (High)42.2
5GPT-5 Mini (Medium)36.3
6Qwen 3 VL 235B A22B (Thinking)30.4
7GPT-5 Mini (Low)30.4
8Qwen 3 VL 32B (Thinking)27.4
9O4 Mini26.7
10GPT-5 Nano22.2
11GPT-4o18.5
12Qwen 3 VL 30B A3B (Thinking)17
13Gemini 2.5 Flash15.6
14Qwen 3 VL 8B (Thinking)14.1
15Qwen 3 VL 8B Instruct13.1

Interactive version: theaggregate.ai/benchmark?slug=visreason-vision-centric-board-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-07.