HoloCount - Joint-Set Aggregation: leaderboard

Metric: Accuracy (%; exact-match accuracy of the single integer count a multimodal model returns zero-shot for an image and a counting question, with a strict integer-only system prompt; the total count of objects from several categories (100 questions); higher is better). Source: arxiv.org. Saturation forecast: Around December 2026. 30 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)82
2Qwen 3.5 122B A10B82
3Qwen 3.5 27B79
4Qwen 3.5 397B A17B78
5Qwen 3.5 35B A3B77
6Gemini 3 Flash (Preview)77
7Qwen 3.5 4B71
8Qwen 3.5 9B69
9Kimi K2.668
10Kimi K2.567
11Gemini 2.5 Pro65
12Claude Opus 4.762
13Qwen 2.5 VL 72B Instruct61
14GPT-5.560
15Claude Opus 4.858

Interactive version: theaggregate.ai/benchmark?slug=holocount-joint-set-aggregation · How It Works · Data refreshed daily, snapshot 2026-09-29.