HoloCount - High-Density Enumeration: leaderboard

Metric: Accuracy (%; exact-match accuracy of the single integer count a multimodal model returns zero-shot for an image and a counting question, with a strict integer-only system prompt; counting densely packed synthetic symbols on a blank canvas (100 questions); higher is better). Source: arxiv.org. Saturation forecast: Around February 2028. 30 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)49
2GPT-5.520
3Qwen 3.5 27B20
4Qwen 3.5 397B A17B14
5Gemini 3 Flash (Preview)12
6Qwen 3.5 122B A10B9
7GPT-5.48
8Qwen 3.5 35B A3B5
9Claude Opus 4.74
10Kimi K2.64
11InternVL3.5-8B4
12Kimi K2.53
13Qwen 3.5 9B3
14Qwen 3 VL 32B Instruct3
15Claude Sonnet 4.52

Interactive version: theaggregate.ai/benchmark?slug=holocount-high-density-enumeration · How It Works · Data refreshed daily, snapshot 2026-09-29.