MMGist - Visual Perception: leaderboard
Metric: Accuracy (%) on the Visual Perception dimension (items kept from BLINK, HallusionBench and RealWorldQA); MMGist: 7,262 curated items from 18 vision-language benchmarks after removing items answerable without the image, items nearly every model solves and items with faulty labels; eight samples per item at temperature 1.0, step-by-step answer in a boxed block, 16,384 output tokens, medium effort where configurable; higher is better. Source: arxiv.org. Saturation forecast: Around May 2028. 27 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) (Medium) | 75.4 |
| 2 | Seed 2.0 Pro | 70.2 |
| 3 | Qwen 3.6 Plus | 67.3 |
| 4 | Seed 2.0 Lite | 65.4 |
| 5 | Seed 2.0 Mini | 63.2 |
| 6 | Gemini 3.1 Flash Lite (Medium) | 62.4 |
| 7 | Qwen 3.6 35B A3B | 61.7 |
| 8 | GPT-5 (Medium) | 58.5 |
| 9 | Qwen 3.5 27B | 58 |
| 10 | Qwen 3.5 9B | 57.6 |
| 11 | Gemma 4 31B | 55.8 |
| 12 | Qwen 3.5 4B | 55.6 |
| 13 | Step3 VL 10B | 53.3 |
| 14 | GPT-5 Mini (Medium) | 52.7 |
| 15 | Gemma 4 26B A4B | 52.1 |
Interactive version: theaggregate.ai/benchmark?slug=mmgist-visual-perception · How It Works · Data refreshed daily, snapshot 2026-09-29.