HoloCount - Differential Comparison: leaderboard

Metric: Accuracy (%; exact-match accuracy of the single integer count a multimodal model returns zero-shot for an image and a counting question, with a strict integer-only system prompt; the difference between the counts of two object sets (100 questions); higher is better). Source: arxiv.org. Saturation forecast: Around December 2026. 30 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)79
2Gemini 3 Flash (Preview)78
3Qwen 3.5 27B76
4Qwen 3.5 397B A17B74
5Kimi K2.670
6Qwen 3.5 122B A10B70
7Qwen 3.5 35B A3B69
8Qwen 3.5 9B67
9Qwen 3.5 4B67
10Kimi K2.566
11Gemini 2.5 Pro65
12Claude Opus 4.859
13GPT-5.558
14Claude Opus 4.758
15MiniMax-M356

Interactive version: theaggregate.ai/benchmark?slug=holocount-differential-comparison · How It Works · Data refreshed daily, snapshot 2026-09-29.