AgentProcessBench - BFCL First-Error Accuracy: leaderboard

Metric: FirstErrAcc (%): share of trajectories where the judge locates the first incorrect step at the same position as the experts (or agrees there is none), on the 250 BFCL trajectories (function calling with the official BFCL tool sets) of AgentProcessBench's 1,000 human-annotated tool-use trajectories (250 per source benchmark, rollouts of five agent models); higher is better. Source: arxiv.org. Saturation forecast: Around May 2028. 20 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.2 Instant58#205
2Kimi K2.5 (Thinking)57.6#139 (Kimi K2.5)
3Kimi K2.5 (Non-reasoning)56.4#139 (Kimi K2.5)
4DeepSeek V3.2 (Thinking)54.4#198 (DeepSeek V3.2)
5GPT-5.2 (Non-reasoning)52.8#105 (GPT-5.2)
6DeepSeek V3.2 (Non-reasoning)50#198 (DeepSeek V3.2)
7GPT-5.2 (Medium)44.4#105 (GPT-5.2)
8Gemini 3 Flash (Preview) (Non-reasoning)40.8#78 (Gemini 3 Flash (Preview))
9Qwen 3 8B (Thinking)38.8#667 (Qwen 3 8B)
10Qwen 3 30B A3B 2507 Instruct36.4#464
11Qwen 3 30B A3B 2507 (Thinking)35.2#366 (Qwen 3 30B A3B 2507)
12Qwen 3 4B 2507 Instruct33.6#745
13Qwen 3 4B 2507 (Thinking)33.6#525 (Qwen 3 4B 2507)
14Qwen 3 4B (Reasoning)33.2#823 (Qwen 3 4B)
15Llama 3.3 70B Instruct30.8#520

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=agentprocessbench-bfcl-first-error-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-11.