PM-Bench: leaderboard

Metric: Set-F1 (%; F1 of the prospective actions the agent selects against the actions actually due, accumulated over the 80 steps and 81 scored tasks of the released simulated week, lures included; single agent with no memory scaffold, heartbeat or subagents; one run per model). Source: arxiv.org. Saturation forecast: Around December 2026. 8 models tracked.

Top models

#ModelScore
1GPT-5.3 Codex78.9
2Mistral-Large-3-675B-Instruct-251275.3
3GPT-5.472.5
4Llama 3.3 70B Instruct64.4
5Mistral Small 3.252.6
6Qwen 3 32B51.4
7Qwen 3 14B43.1
8Qwen 3 8B42

Interactive version: theaggregate.ai/benchmark?slug=pm-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.