VLM-DeflectionBench - Realistic Accuracy: leaderboard

Metric: Accuracy (%) of knowledge-based visual question answering in the Realistic retrieval scenario (two gold and two negative contexts mixed), judged by GPT-4o with the SimpleQA protocol; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 19 models tracked.

Top models

#ModelScore
1GPT-559.5
2Gemini 2.5 Pro51
3InternVL3-38B50.3
4Gemini 2.5 Flash47
5Qwen 2.5 VL 32B Instruct45.2
6Gemma 3 27B42.9
7Pixtral-12B42.6
8GLM-4.1V-9B (Thinking)37.4
9Claude Sonnet 433.3
10Claude Opus 432.1
11Mistral Small 3.123.5

Interactive version: theaggregate.ai/benchmark?slug=vlm-deflectionbench-realistic-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-07.