AgentProcessBench - First-Error Accuracy: leaderboard

Metric: FirstErrAcc (%): share of trajectories where the judge locates the first incorrect step at the same position as the experts (or agrees there is none), averaged over the four source subsets of AgentProcessBench's 1,000 human-annotated tool-use trajectories (250 per source benchmark, rollouts of five agent models); higher is better. Source: arxiv.org. Saturation forecast: Around August 2028. 20 models tracked.

Top models

#ModelScoreOverall rank
1Kimi K2.5 (Thinking)62.4#139 (Kimi K2.5)
2GPT-5.2 Instant61.1#205
3DeepSeek V3.2 (Thinking)59.6#198 (DeepSeek V3.2)
4GPT-5.2 (Non-reasoning)58.3#105 (GPT-5.2)
5GPT-5.2 (Medium)57.5#105 (GPT-5.2)
6Kimi K2.5 (Non-reasoning)56.8#139 (Kimi K2.5)
7DeepSeek V3.2 (Non-reasoning)55.2#198 (DeepSeek V3.2)
8Gemini 3 Flash (Preview) (Non-reasoning)55#78 (Gemini 3 Flash (Preview))
9Qwen 3 30B A3B 2507 (Thinking)52#366 (Qwen 3 30B A3B 2507)
10Qwen 3 8B (Thinking)46#667 (Qwen 3 8B)
11Qwen 3 4B 2507 Instruct44.4#745
12Qwen 3 4B 2507 (Thinking)44.4#525 (Qwen 3 4B 2507)
13Qwen 3 4B (Reasoning)44.1#823 (Qwen 3 4B)
14Qwen 3 30B A3B 2507 Instruct43.9#464
15Llama 3.3 70B Instruct42.1#520

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=agentprocessbench-first-error-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-11.