HoloCount - Categorical Cardinality: leaderboard

Metric: Accuracy (%; exact-match accuracy of the single integer count a multimodal model returns zero-shot for an image and a counting question, with a strict integer-only system prompt; counting the number of distinct object kinds present, such as kinds of fruit (101 questions); higher is better). Source: arxiv.org. Saturation forecast: Around December 2026. 30 models tracked.

Top models

#ModelScore
1Kimi K2.689.1
2Claude Opus 4.784.2
3Gemini 2.5 Pro82.2
4Kimi K2.582.2
5Claude Opus 4.882.2
6Qwen 3.5 35B A3B82.2
7Qwen 3.5 122B A10B82.2
8GPT-5.581.2
9Qwen 3.5 9B81.2
10Gemini 3 Flash (Preview)81.2
11Qwen 3.5 397B A17B80.2
12Qwen 3.5 27B80.2
13Qwen 3.5 4B80.2
14GPT-4.177.2
15MiniMax-M375.2

Interactive version: theaggregate.ai/benchmark?slug=holocount-categorical-cardinality · How It Works · Data refreshed daily, snapshot 2026-09-29.