Blind-Spots-Bench - Multimodal-to-Text: leaderboard

Metric: Accuracy (%; mean@4 over four samples on the questions with image input and text output, vision-language models only; 235 human-authored questions with reference solutions graded correct or incorrect by gemini-3-flash with code execution; no tools; thinking enabled at medium effort where available; max 32,768 output tokens). Source: arxiv.org. Saturation forecast: Around December 2026. 22 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview) (Medium)66.9
2Gemini 3 Flash (Preview) (Medium)64.5
3GPT-5.558.7
4GPT-5.4 (Medium)58.1
5GPT-5.2 (Medium)50
6Qwen 3.5 122B A10B48.8
7Qwen 3.5 397B A17B48.3
8Qwen 3.5 35B A3B47.1
9GPT-5.4 Mini (Medium)45.3
10Kimi K2.542.4
11Kimi K2.641.9
12gemma-4-26B-A4B-it (Thinking)40.7
13GPT-537.8
14Gemini 2.5 Pro36
15GPT-5 Mini33.1

Interactive version: theaggregate.ai/benchmark?slug=blind-spots-bench-multimodal-to-text · How It Works · Data refreshed daily, snapshot 2026-09-29.