MP-Bench (Failure Attribution) - Hand-Crafted MAS: leaderboard

Metric: nDCG@5 with exponential gain (0-1, times 100) of the model's ranking of failure-inducing steps against the ranking by annotator consensus, on MP-Bench's 169 failed hand-crafted (Magentic-One) multi-agent executions from GAIA and AssistantBench, three expert annotators per execution log; the model runs All-at-Once attribution zero-shot 3 times at temperature 1.0 and its runs are consolidated by consensus rate; higher is better. Source: arxiv.org. Saturation forecast: Around August 2028. 6 models tracked.

Top models

#ModelScoreOverall rank
1Claude Sonnet 4.543.97#138
2O3 Mini43.67#266
3GPT-4.143.13#240
4GPT-OSS-120B42.45#330
5GPT-5.137.47#131
6Qwen 3 8B29.44#667

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=mp-bench-failure-attribution-hand-crafted-mas · How It Works · Data refreshed daily, snapshot 2026-10-11.