ComplexFuncBench: leaderboard

1,000 multi-step function-calling tasks over 43 live travel APIs (hotels, flights, attractions, car rental, taxi) with 128k contexts, from Zhipu AI and Tsinghua (2025); success rate, GPT-4o 60.5%.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturated. 20 models tracked.

Top models

#ModelScore
1GPT-4o (2024-11-20)66.5
2GPT-4.165.5
3GPT-4.563
4Claude 3.5 Sonnet61
5GPT-4o (2024-08-06)60.5
6GPT-4 Turbo49.5
7GPT-4.1 Mini49.3
8O1 (High)47.6
9Claude 3.5 Haiku45.8
10Qwen 2.5 72B40.1
11GPT-4o Mini38.6
12Mistral Large 2 (Jul)20.1
13O3 Mini (High)17.6
14Llama 3.1 8B10
15glm-4-9B9.4

Interactive version: theaggregate.ai/benchmark?slug=complexfuncbench · How It Works · Data refreshed daily, snapshot 2026-09-05.