HoloCount - Scale Attribution: leaderboard

Metric: Accuracy (%; exact-match accuracy of the single integer count a multimodal model returns zero-shot for an image and a counting question, with a strict integer-only system prompt; counting objects of a stated size (100 questions); higher is better). Source: arxiv.org. Saturation forecast: Around May 2027. 30 models tracked.

Top models

#ModelScore
1Qwen 3.5 27B86
2Gemini 3.1 Pro (Preview)85
3Qwen 3.5 9B85
4Gemini 3 Flash (Preview)85
5Kimi K2.684
6Qwen 3.5 122B A10B84
7Qwen 3.5 35B A3B83
8Kimi K2.582
9Qwen 3.5 397B A17B82
10Qwen 2.5 VL 72B Instruct81
11Qwen 3.5 4B78
12Claude Opus 4.877
13Qwen 3 VL 32B Instruct77
14Gemini 2.5 Pro74
15GPT-5.573

Interactive version: theaggregate.ai/benchmark?slug=holocount-scale-attribution · How It Works · Data refreshed daily, snapshot 2026-09-29.