GroupToM-Bench - Collective Outcome Prediction: leaderboard

Metric: strict exact-match accuracy (%) over multi-select options (random 6.7) on Level 6 collective outcome prediction (non-linear group outcomes such as the Abilene paradox), multiple choice; 240 multimodal multi-agent scenarios (scene image plus dialogue); higher is better. Source: arxiv.org. Saturation forecast: Around August 2027. 11 models tracked.

Top models

#ModelScore
1GPT-4o48.6
2Gemini 3 Pro48.3
3GPT-545
4Claude Haiku 4.544.1
5GPT-5 Mini42.5
6GPT-5 Nano32.5
7Qwen 2.5 VL 7B Instruct31.7
8InternVL3.5-8B26.2
9Qwen 2 VL 7B17.2

Interactive version: theaggregate.ai/benchmark?slug=grouptom-bench-collective-outcome-prediction · How It Works · Data refreshed daily, snapshot 2026-09-29.