PRISM-VLM - Visual Discrimination: leaderboard

Metric: Per-axis mean score (%; forced A/B choice between the gold caption and a one-detail minimal edit (chance 50%); over PRISM-VLM's 6,238 items recycled from 15 public VLM benchmarks (five seeds of 100 items per benchmark); GPT-5 (low effort) synthesizes the perturbations and grades the open-ended axes; each model at the lowest reasoning effort its provider exposes, temperature 0). Source: arxiv.org. Saturation forecast: Around December 2026. 42 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Minimal)93.7
2Gemini 2.5 Flash (Non-reasoning)91.5
3Gemini 2.0 Flash90
4Gemini 2.0 Flash Lite87.6
5Gemini 2.5 Flash Lite87.5
6GPT-4.1 Mini86.7
7Molmo2-8B86.4
8GPT-5 Mini (Minimal)86
9Qwen 3.5 9B (Non-reasoning)85.9
10Qwen 3 VL 8B Instruct85.8
11Gemma 4 E4B84
12Claude Haiku 4.583.9
13Qwen 3 VL 4B Instruct83
14GPT-5.4 Mini82.7
15Grok 4 Fast (Non-reasoning)82.5

Interactive version: theaggregate.ai/benchmark?slug=prism-vlm-visual-discrimination · How It Works · Data refreshed daily, snapshot 2026-09-26.