HoloCount - Complementary Exclusion: leaderboard

Metric: Accuracy (%; exact-match accuracy of the single integer count a multimodal model returns zero-shot for an image and a counting question, with a strict integer-only system prompt; counting objects in one set but not another (101 questions); higher is better). Source: arxiv.org. Saturation forecast: Around January 2027. 30 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)84.2
2Qwen 3.5 397B A17B80.2
3Gemini 3 Flash (Preview)80.2
4Qwen 3.5 122B A10B79.2
5Qwen 3.5 35B A3B78.2
6Qwen 3.5 27B76.2
7Gemini 2.5 Pro74.3
8Kimi K2.674.3
9Qwen 3.5 4B70.3
10Qwen 3.5 9B67.3
11GPT-5.566.3
12Kimi K2.563.4
13Claude Opus 4.863.4
14GPT-5.461.4
15GPT-4.160.4

Interactive version: theaggregate.ai/benchmark?slug=holocount-complementary-exclusion · How It Works · Data refreshed daily, snapshot 2026-09-29.