AgentBench FC — leaderboard

Function-calling agent benchmark across ALFWorld, database, knowledge-graph, operating-system, and WebShop environments, requiring tool calls that complete multi-step tasks.

Metric: Avg Success Rate (%). Source: github.com. Status: saturation imminent. 25 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.558.9
2Claude Sonnet 4.5 (Thinking)58.3
3Claude Sonnet 4 (Thinking)58.2
4Claude Sonnet 457.4
5Claude 3.7 Sonnet53.2
6GPT-552.2
7Claude 3.7 Sonnet (Thinking)50
8DeepSeek R149.3
9O3 Mini40.9
10Qwen 2.5 72B Instruct40.8
11O4 Mini39.7
12GPT-4o39.6
13Qwen 2.5 32B Instruct37.2
14DeepSeek V336.1
15Qwen 2.5 14B Instruct27.2

Interactive version: theaggregate.ai/benchmark?slug=agentbench-fc · How the rankings work · Data refreshed daily, snapshot 2026-07-22.