AgentProcessBench - First-Error Accuracy: leaderboard
Metric: FirstErrAcc (%): share of trajectories where the judge locates the first incorrect step at the same position as the experts (or agrees there is none), averaged over the four source subsets of AgentProcessBench's 1,000 human-annotated tool-use trajectories (250 per source benchmark, rollouts of five agent models); higher is better. Source: arxiv.org. Saturation forecast: Around August 2028. 20 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Kimi K2.5 (Thinking) | 62.4 | #139 (Kimi K2.5) |
| 2 | GPT-5.2 Instant | 61.1 | #205 |
| 3 | DeepSeek V3.2 (Thinking) | 59.6 | #198 (DeepSeek V3.2) |
| 4 | GPT-5.2 (Non-reasoning) | 58.3 | #105 (GPT-5.2) |
| 5 | GPT-5.2 (Medium) | 57.5 | #105 (GPT-5.2) |
| 6 | Kimi K2.5 (Non-reasoning) | 56.8 | #139 (Kimi K2.5) |
| 7 | DeepSeek V3.2 (Non-reasoning) | 55.2 | #198 (DeepSeek V3.2) |
| 8 | Gemini 3 Flash (Preview) (Non-reasoning) | 55 | #78 (Gemini 3 Flash (Preview)) |
| 9 | Qwen 3 30B A3B 2507 (Thinking) | 52 | #366 (Qwen 3 30B A3B 2507) |
| 10 | Qwen 3 8B (Thinking) | 46 | #667 (Qwen 3 8B) |
| 11 | Qwen 3 4B 2507 Instruct | 44.4 | #745 |
| 12 | Qwen 3 4B 2507 (Thinking) | 44.4 | #525 (Qwen 3 4B 2507) |
| 13 | Qwen 3 4B (Reasoning) | 44.1 | #823 (Qwen 3 4B) |
| 14 | Qwen 3 30B A3B 2507 Instruct | 43.9 | #464 |
| 15 | Llama 3.3 70B Instruct | 42.1 | #520 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=agentprocessbench-first-error-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-11.