PM-Bench (Hierarchical Multi-Agent): leaderboard

Metric: Set-F1 (%; F1 of the prospective actions the agent selects against the actions actually due, accumulated over the 80 steps and 81 scored tasks of the released simulated week, lures included; coordinator with three specialist subagents whose suggested state queries are all executed; one run per model). Source: arxiv.org. Saturation forecast: Around March 2027. 8 models tracked.

Top models

#ModelScore
1Mistral-Large-3-675B-Instruct-251258.1
2GPT-5.456.5
3GPT-5.3 Codex55.9
4Llama 3.3 70B Instruct50.9
5Qwen 3 14B43.6
6Qwen 3 32B37.5
7Mistral Small 3.233.6
8Qwen 3 8B25.7

Interactive version: theaggregate.ai/benchmark?slug=pm-bench-hierarchical-multi-agent · How It Works · Data refreshed daily, snapshot 2026-09-29.