Smolagents Leaderboard — leaderboard
Smolagents leaderboard comparing agent action styles on SimpleQA, GAIA, and MATH tasks using public Hugging Face result snapshots.
Metric: Average Accuracy (%). Source: huggingface.co. Status: saturated. 31 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4.5 (Preview) | 78.75 |
| 2 | Claude 3.7 Sonnet (20250219) | 77.33 |
| 3 | O1 | 76.67 |
| 4 | O3 Mini | 74.96 |
| 5 | Claude 3.5 Sonnet | 74.75 |
| 6 | DeepSeek R1 | 72.67 |
Interactive version: theaggregate.ai/benchmark?slug=smolagents-leaderboard · How the rankings work · Data refreshed daily, snapshot 2026-07-22.