VIA-Bench (LLM Judge) - General Visual Illusions: leaderboard

Metric: Accuracy (%) on the general visual illusions questions, answer letter read from the response by a GPT-4.1-mini judge; VIA-Bench's 1,004 human-verified multiple-choice questions on illusory and anomalous images, every option set including a never-correct 'Not Sure' option; zero-shot with an answer-format instruction, temperature 0.8 where supported (model defaults otherwise), mean of five runs; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 24 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Pro (Preview)67.91#64
2O367.75#121
3Gemini 2.5 Pro65.74#145
4GLM-4.5V65.74#339
5Qwen 3 VL 235B A22B (Thinking)65.58#228 (Qwen 3 VL 235B A22B)
6GPT-4o (2024-11-20)64.19#369
7Qwen 2.5 VL 72B Instruct62.79#364
8GPT-5 Mini61.86#176
9O4 Mini61.24#172
10Qwen 3 VL 30B A3B (Thinking)60.78#338 (Qwen 3 VL 30B A3B)
11GPT-4o ChatGPT54.88
12Claude Sonnet 4 (20250514)53.95#211
13Qwen 2.5 VL 7B Instruct53.33#643
14Claude Opus 4.1 (20250805)48.84#126
15GPT-5 Chat48.53#233

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=via-bench-llm-judge-general-visual-illusions · How It Works · Data refreshed daily, snapshot 2026-10-11.