Smolagents Leaderboard — leaderboard

Smolagents leaderboard comparing agent action styles on SimpleQA, GAIA, and MATH tasks using public Hugging Face result snapshots.

Metric: Average Accuracy (%). Source: huggingface.co. Status: saturated. 31 models tracked.

Top models

#ModelScore
1GPT-4.5 (Preview)78.75
2Claude 3.7 Sonnet (20250219)77.33
3O176.67
4O3 Mini74.96
5Claude 3.5 Sonnet74.75
6DeepSeek R172.67

Interactive version: theaggregate.ai/benchmark?slug=smolagents-leaderboard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.