ComplexFuncBench — leaderboard

ComplexFuncBench evaluates model capability on tool use tasks from the linked upstream source with Score as the primary reported metric.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 12 models tracked.

Top models

#ModelScore
1Claude 3.5 Sonnet61
2GPT-4o (2024-08-06)60.5
3GPT-4 Turbo49.5
4Claude 3.5 Haiku45.8
5Qwen 2.5 72B Instruct40.1
6Mistral Large 2 (Nov) Instruct (2411)20.1
7glm-4-9B9.4
8Qwen 2.5 7B5
9Llama 3.1 405B Instruct4
10Llama 3.1 70B2.7
11Llama 3.1 8B0.1

Interactive version: theaggregate.ai/benchmark?slug=complexfuncbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.