HoloCount: leaderboard

Metric: Overall accuracy (%; exact-match accuracy of the single integer count a multimodal model returns zero-shot for an image and a counting question, with a strict integer-only system prompt, over all 2,480 questions in the 20 subsets; higher is better). Source: arxiv.org. Saturation forecast: Around January 2028. 30 models tracked.

Top models

#ModelScore
1Qwen 3.5 397B A17B76.9
2Qwen 3.5 27B75.9
3Qwen 3.5 122B A10B75.4
4Qwen 3.5 35B A3B75
5Gemini 3 Flash (Preview)74.8
6Gemini 3.1 Pro (Preview)74.7
7Kimi K2.673
8Qwen 3.5 9B71.4
9Kimi K2.568.8
10Qwen 3.5 4B68.6
11Gemini 2.5 Pro67.9
12GPT-5.567.6
13Claude Opus 4.762.8
14Qwen 3 VL 32B Instruct62.7
15Claude Opus 4.862.6

Interactive version: theaggregate.ai/benchmark?slug=holocount · How It Works · Data refreshed daily, snapshot 2026-09-29.