GroupToM-Bench - Mechanistic Attribution: leaderboard

Metric: mean judge score (0-100) on Level 7 mechanistic attribution (structural explanation of a collective failure), open-ended, scored 0 to 100 by a GPT-5 judge against expert references; 240 multimodal multi-agent scenarios (scene image plus dialogue); higher is better. Source: arxiv.org. Saturation forecast: Around August 2027. 11 models tracked.

Top models

#ModelScore
1Gemini 3 Pro64.2
2GPT-561
3GPT-5 Mini59.9
4GPT-4o53.4
5Claude Haiku 4.552.9
6GPT-5 Nano49
7InternVL3.5-8B47.5
8Qwen 2.5 VL 7B Instruct43.4
9Qwen 2 VL 7B41.3

Interactive version: theaggregate.ai/benchmark?slug=grouptom-bench-mechanistic-attribution · How It Works · Data refreshed daily, snapshot 2026-09-29.