ProLLM - Function Calling — leaderboard

ProLLM benchmark evaluating LLM function calling accuracy across practical API invocation tasks.

Metric: Score (%). Source: www.prollm.ai. Status: saturation imminent. 90 models tracked.

Top models

#ModelScore
1O197.3
2Gemini 2.0 Flash96.7
3GPT-4o Mini92.3
4Grok 3 Mini92.3
5GPT-4.591.9
6Grok Beta91.8
7Grok 391.1
8O3 Mini (Medium)91.1
9O3 Mini (High)91
10DeepSeek R190.4
11O1 Preview90
12GPT-4.189.9
13GPT-4.1 Mini89.8
14DeepSeek V389.8
15Mistral Small 389.6

Interactive version: theaggregate.ai/benchmark?slug=prollm-function-calling · How the rankings work · Data refreshed daily, snapshot 2026-07-22.