GAIA — leaderboard

General AI Assistants benchmark - 466 questions requiring multi-step reasoning, web browsing, and tool use. Questions are easy for humans (~92%) but hard for AI. Three difficulty levels.

Metric: Accuracy (%). Source: huggingface.co. Status: saturation imminent. 3460 models tracked.

Top models

#ModelScore
1GPT-5.277.41
2DeepSeek V3.276.74
3GPT-566.45
4Gemini 3.1 Pro (Preview)55.48
5Gemini 3 Flash41.53
6GPT-4o (2024-08-06)28.9
7DeepSeek V4 Flash23.59
8GPT-5 Nano23.26
9GPT-5.119.6
10GPT-4 Turbo11.3
11Gemma 4 26B A4B (IT)8.31
12GPT-4o Mini (2024-07-18)4.65
13GPT-43.99
14GPT-3.52.66
15Qwen 2.5 7B0.66

Interactive version: theaggregate.ai/benchmark?slug=gaia · How the rankings work · Data refreshed daily, snapshot 2026-07-22.