REAL Evals — leaderboard
Realistic Evaluations for Agents in the Loop: real-world task completion benchmark with verified model performance on practical tasks.
Metric: Task Completion (%). Source: www.realevals.xyz. Status: saturation imminent. 33 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Flash | 30 |
| 2 | Claude Sonnet 4.5 | 29.2 |
| 3 | Gemini 2.5 Pro | 25.3 |
| 4 | Claude Sonnet 4 | 24 |
| 5 | Claude 3.7 Sonnet (Thinking) | 24 |
| 6 | GPT-5 | 23.6 |
| 7 | Claude Opus 4 (Thinking) | 19.7 |
| 8 | Grok 4 Fast | 19.3 |
| 9 | Claude Sonnet 4 (Thinking) | 18.9 |
| 10 | O3 | 16.7 |
| 11 | Claude 3.7 Sonnet | 16.3 |
| 12 | DeepSeek V3.2 Exp | 12.4 |
| 13 | O3 Mini | 12 |
| 14 | GPT-4.1 | 11.6 |
| 15 | GPT-5 Nano | 9.9 |
Interactive version: theaggregate.ai/benchmark?slug=real-evals · How the rankings work · Data refreshed daily, snapshot 2026-07-22.