UNICBench (Image): leaderboard

Metric: Exact-count hit rate (%): share of parsed numeric answers equal to the annotated count over UNICBench's 5,508 image counting questions (5,300 images; everyday objects, remote sensing, crowds, cells), temperature 0 (1.0 where a model fixes it), 4,096-token cap, GPT-5 and GPT-5-mini at minimal and o3 at low reasoning effort; unparsable responses are left out of the rate; higher is better. Source: arxiv.org. Saturation forecast: Around 2032. 21 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro20.3#145
2O3 (Low)18.5#121 (O3)
3Qwen 3 VL 30B A3B (Thinking)18.3#338 (Qwen 3 VL 30B A3B)
4InternVL3-78B17.8#345
5Qwen 2.5 VL 72B Instruct17.4#364
6Qwen 3 VL 30B A3B Instruct17.3#365
7GPT-5 Mini (Minimal)17.2#176 (GPT-5 Mini)
8O4 Mini17.1#172
9GPT-5 (Minimal)16.8#91 (GPT-5)
10GPT-4o16.6#333
11Qwen 2.5 VL 7B Instruct16.6#643
12GLM-4.5V16#339
13GLM-4.1V-9B (Thinking)15.7#457 (GLM-4.1V-9B)
14Gemini 2.5 Flash15.6#237
15Claude Sonnet 4 (20250514)14.9#211

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=unicbench-image · How It Works · Data refreshed daily, snapshot 2026-10-11.