PRISM-VLM - Multi-Question Robustness: leaderboard

Metric: Per-axis mean score (%; share of items whose 3-5 bundled questions about one image are all answered correctly; over PRISM-VLM's 6,238 items recycled from 15 public VLM benchmarks (five seeds of 100 items per benchmark); GPT-5 (low effort) synthesizes the perturbations and grades the open-ended axes; each model at the lowest reasoning effort its provider exposes, temperature 0). Source: arxiv.org. Saturation forecast: Around 2029. 42 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Minimal)56
2GPT-4.1 Mini50.7
3Qwen 3.5 9B (Non-reasoning)50.5
4GPT-5 Mini (Minimal)49.4
5Gemini 2.5 Flash (Non-reasoning)48.7
6GPT-5.4 Mini48.6
7Qwen 3.5 4B (Non-reasoning)48.5
8Gemini 2.0 Flash46.8
9Qwen 3 VL 8B Instruct44.4
10Gemini 2.0 Flash Lite43.7
11Qwen 3 VL 4B Instruct41.5
12Gemini 2.5 Flash Lite40.8
13Molmo2-8B39.7
14Nova 2 Lite38.7
15Claude Haiku 4.537.4

Interactive version: theaggregate.ai/benchmark?slug=prism-vlm-multi-question-robustness · How It Works · Data refreshed daily, snapshot 2026-09-26.