VLM-DeflectionBench - Oracle Accuracy: leaderboard

Metric: Accuracy (%) of knowledge-based visual question answering in the Oracle retrieval scenario (gold evidence only), judged by GPT-4o with the SimpleQA protocol; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 19 models tracked.

Top models

#ModelScore
1GPT-573.1
2InternVL3-38B64.2
3Pixtral-12B62.7
4Qwen 2.5 VL 32B Instruct61
5Gemini 2.5 Pro59.8
6Gemma 3 27B59.5
7Gemini 2.5 Flash58.8
8GLM-4.1V-9B (Thinking)55.2
9Claude Opus 449.1
10Claude Sonnet 447.5
11Mistral Small 3.142.6

Interactive version: theaggregate.ai/benchmark?slug=vlm-deflectionbench-oracle-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-07.