HoloCount - General Semantic Attribution: leaderboard

Metric: Accuracy (%; exact-match accuracy of the single integer count a multimodal model returns zero-shot for an image and a counting question, with a strict integer-only system prompt; counting objects with a general semantic attribute (101 questions); higher is better). Source: arxiv.org. Saturation forecast: Around March 2027. 30 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)87.1
2Gemini 3 Flash (Preview)86.1
3Qwen 3.5 35B A3B83.2
4Qwen 3.5 27B83.2
5Kimi K2.682.2
6Qwen 3.5 397B A17B80.2
7Qwen 3.5 122B A10B77.2
8Qwen 3.5 4B76.2
9Qwen 3.5 9B75.2
10Gemini 2.5 Pro71.3
11GPT-5.568.3
12Kimi K2.568.3
13Claude Opus 4.868.3
14Qwen 2.5 VL 32B Instruct67.3
15MiniMax-M366.3

Interactive version: theaggregate.ai/benchmark?slug=holocount-general-semantic-attribution · How It Works · Data refreshed daily, snapshot 2026-09-29.