Gorilla API Bench (BFCL) — leaderboard

Metric: Overall Accuracy (%). Source: gorilla.cs.berkeley.edu. 109 models tracked.

Top models

#ModelScore
1Claude Opus 4.5 (20251101)77.47
2Claude Sonnet 4.573.24
3Gemini 3 Pro (Preview)72.51
4GLM-4.6 (Thinking)72.38
5Grok 4.1 Fast (Reasoning)69.57
6Claude Haiku 4.5 (20251001)68.7
7O3 (2025-04-16)63.05
8Grok 462.97
9Kimi K259.06
10Grok 4.1 Fast (Non-reasoning)58.29
11DeepSeek V3.2 Exp (Thinking)56.73
12Gemini 2.5 Flash56.24
13GPT-5.255.87
14GPT-5 Mini55.46
15DeepSeek V3.2 Exp54.12

Interactive version: theaggregate.ai/benchmark?slug=gorilla-api-bench-bfcl · How the rankings work · Data refreshed daily, snapshot 2026-07-22.