AgentProcessBench - BFCL First-Error Accuracy: leaderboard
Metric: FirstErrAcc (%): share of trajectories where the judge locates the first incorrect step at the same position as the experts (or agrees there is none), on the 250 BFCL trajectories (function calling with the official BFCL tool sets) of AgentProcessBench's 1,000 human-annotated tool-use trajectories (250 per source benchmark, rollouts of five agent models); higher is better. Source: arxiv.org. Saturation forecast: Around May 2028. 20 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT-5.2 Instant | 58 | #205 |
| 2 | Kimi K2.5 (Thinking) | 57.6 | #139 (Kimi K2.5) |
| 3 | Kimi K2.5 (Non-reasoning) | 56.4 | #139 (Kimi K2.5) |
| 4 | DeepSeek V3.2 (Thinking) | 54.4 | #198 (DeepSeek V3.2) |
| 5 | GPT-5.2 (Non-reasoning) | 52.8 | #105 (GPT-5.2) |
| 6 | DeepSeek V3.2 (Non-reasoning) | 50 | #198 (DeepSeek V3.2) |
| 7 | GPT-5.2 (Medium) | 44.4 | #105 (GPT-5.2) |
| 8 | Gemini 3 Flash (Preview) (Non-reasoning) | 40.8 | #78 (Gemini 3 Flash (Preview)) |
| 9 | Qwen 3 8B (Thinking) | 38.8 | #667 (Qwen 3 8B) |
| 10 | Qwen 3 30B A3B 2507 Instruct | 36.4 | #464 |
| 11 | Qwen 3 30B A3B 2507 (Thinking) | 35.2 | #366 (Qwen 3 30B A3B 2507) |
| 12 | Qwen 3 4B 2507 Instruct | 33.6 | #745 |
| 13 | Qwen 3 4B 2507 (Thinking) | 33.6 | #525 (Qwen 3 4B 2507) |
| 14 | Qwen 3 4B (Reasoning) | 33.2 | #823 (Qwen 3 4B) |
| 15 | Llama 3.3 70B Instruct | 30.8 | #520 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=agentprocessbench-bfcl-first-error-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-11.