ACEBench: leaderboard

2,000 tool-calling test cases in parallel English and Chinese over 4,538 APIs, split into Normal, Special (ambiguous instructions) and multi-turn Agent sets (Huawei Noah's Ark Lab, 2025); accuracy.

Metric: Overall (self-reported). Source: benchmarklist.com. Status: saturated. 29 models tracked.

Top models

#ModelScore
1GPT-4o (2024-11-20)89.6
2GPT-4 Turbo88.6
3Qwen Max81.7
4O1 Preview80.6
5DeepSeek V3 Chat78.5
6GPT-4o Mini (2024-07-18)76
7Claude 3.5 Sonnet (20241022)75.6
8Gemini 1.5 Pro72.8
9O1 Mini72.2

Interactive version: theaggregate.ai/benchmark?slug=acebench · How It Works · Data refreshed daily, snapshot 2026-09-05.