BunchCount - Group Counting: leaderboard
Metric: Mean absolute error (semantic groups per image, lower is better; group-level test split of 310 real images with 14 categories unseen in training, e.g. bunches of grapes, stacks of plates, pairs of shoes; vision-language models zero-shot with reasoning disabled, temperature 0, one-integer answer). Source: arxiv.org. Saturation forecast: Around August 2027. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.6 Sol (Non-reasoning) | 5.46 |
| 2 | Qwen 3.7 Max (Non-reasoning) | 5.55 |
| 3 | Qwen 3.5 27B (Non-reasoning) | 5.84 |
| 4 | Qwen 3 VL 8B Instruct | 10.67 |
Interactive version: theaggregate.ai/benchmark?slug=bunchcount-group-counting · How It Works · Data refreshed daily, snapshot 2026-09-26.