ACEBench — leaderboard

ACEBench evaluates model capability on tool use tasks from the linked upstream source with Overall as the primary reported metric.

Metric: Overall (self-reported). Source: benchmarklist.com. Status: saturation imminent. 29 models tracked.

Top models

#ModelScore
1GPT-4o (2024-11-20)89.6
2GPT-4 Turbo88.6
3O1 Preview80.6
4DeepSeek V378.5
5GPT-4o Mini (2024-07-18)76
6Claude 3.5 Sonnet75.6
7Gemini 1.5 Pro72.8
8O1 Mini72.2

Interactive version: theaggregate.ai/benchmark?slug=acebench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.