REAL Evals — leaderboard

Realistic Evaluations for Agents in the Loop: real-world task completion benchmark with verified model performance on practical tasks.

Metric: Task Completion (%). Source: www.realevals.xyz. Status: saturation imminent. 33 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash30
2Claude Sonnet 4.529.2
3Gemini 2.5 Pro25.3
4Claude Sonnet 424
5Claude 3.7 Sonnet (Thinking)24
6GPT-523.6
7Claude Opus 4 (Thinking)19.7
8Grok 4 Fast19.3
9Claude Sonnet 4 (Thinking)18.9
10O316.7
11Claude 3.7 Sonnet16.3
12DeepSeek V3.2 Exp12.4
13O3 Mini12
14GPT-4.111.6
15GPT-5 Nano9.9

Interactive version: theaggregate.ai/benchmark?slug=real-evals · How the rankings work · Data refreshed daily, snapshot 2026-07-22.