HoloCount - Small-Scale Enumeration: leaderboard

Metric: Accuracy (%; exact-match accuracy of the single integer count a multimodal model returns zero-shot for an image and a counting question, with a strict integer-only system prompt; counting objects smaller than 32 by 32 pixels (101 questions); higher is better). Source: arxiv.org. Saturation forecast: Around 2029. 30 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Preview)74.3
2Gemini 3.1 Pro (Preview)73.3
3Qwen 3.5 397B A17B60.4
4Qwen 3.5 35B A3B60.4
5Qwen 3.5 122B A10B58.4
6Qwen 3.5 27B52.5
7Qwen 3.5 9B50.5
8Qwen 2.5 VL 72B Instruct43.6
9GPT-5.542.6
10Qwen 3.5 4B42.6
11Claude Opus 4.740.6
12Kimi K2.640.6
13Gemini 2.5 Pro39.6
14Claude Opus 4.838.6
15Kimi K2.537.6

Interactive version: theaggregate.ai/benchmark?slug=holocount-small-scale-enumeration · How It Works · Data refreshed daily, snapshot 2026-09-29.