PRISM-VLM - Counting: leaderboard

Metric: Per-axis mean score (%; exact-integer accuracy on synthesized counting questions; over PRISM-VLM's 6,238 items recycled from 15 public VLM benchmarks (five seeds of 100 items per benchmark); GPT-5 (low effort) synthesizes the perturbations and grades the open-ended axes; each model at the lowest reasoning effort its provider exposes, temperature 0). Source: arxiv.org. Saturation forecast: Around 2030. 42 models tracked.

Top models

#ModelScore
1Qwen 3.5 9B (Non-reasoning)66.4
2Gemini 3 Flash (Minimal)65.1
3Qwen 3.5 4B (Non-reasoning)64.6
4GPT-4.1 Mini61.8
5Qwen 3 VL 8B Instruct61.8
6GPT-5.4 Mini61.7
7Gemini 2.5 Flash (Non-reasoning)61.7
8GPT-5 Mini (Minimal)61.5
9Qwen 3 VL 4B Instruct58.5
10Qwen 3.5 2B (Non-reasoning)57.9
11Gemini 2.0 Flash57.3
12Nova 2 Lite56.2
13Gemini 2.0 Flash Lite55.1
14Gemma 4 E4B54.8
15Ministral-3-8B-Instruct-251254.5

Interactive version: theaggregate.ai/benchmark?slug=prism-vlm-counting · How It Works · Data refreshed daily, snapshot 2026-09-26.