BunchCount - Individual Counting: leaderboard
Metric: Mean absolute error (individual instances per image, lower is better; individual-level test split, the 310 real images of 14 categories unseen in training, about 90 instances per image; vision-language models zero-shot with reasoning disabled, temperature 0, one-integer answer). Source: arxiv.org. Saturation forecast: Around July 2027. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.6 Sol (Non-reasoning) | 15.94 |
| 2 | Qwen 3.7 Max (Non-reasoning) | 31.02 |
| 3 | Qwen 3.5 27B (Non-reasoning) | 38.28 |
| 4 | Qwen 3 VL 8B Instruct | 40.07 |
| 5 | GLM-4.5V (Non-reasoning) | 44.5 |
| 6 | InternVL3.5-8B | 46.56 |
Interactive version: theaggregate.ai/benchmark?slug=bunchcount-individual-counting · How It Works · Data refreshed daily, snapshot 2026-09-29.