PM-Bench (Auto Heartbeat 30m): leaderboard

Metric: Set-F1 (%; F1 of the prospective actions the agent selects against the actions actually due, accumulated over the 80 steps and 81 scored tasks of the released simulated week, lures included; single agent with a heartbeat switched on every 30 virtual minutes; one run per model). Source: arxiv.org. Saturation forecast: Around December 2026. 8 models tracked.

Top models

#ModelScore
1Mistral-Large-3-675B-Instruct-251274.7
2GPT-5.474.1
3GPT-5.3 Codex71
4Llama 3.3 70B Instruct63.4
5Qwen 3 32B57.5
6Mistral Small 3.251.6
7Qwen 3 8B39.8
8Qwen 3 14B30.3

Interactive version: theaggregate.ai/benchmark?slug=pm-bench-auto-heartbeat-30m · How It Works · Data refreshed daily, snapshot 2026-09-29.