AGI-Eval Community - Cognition: leaderboard

Metric: Accuracy (%). Source: agi-eval.cn. 141 models tracked.

Top models

#ModelScore
1Claude Fable 588.43
2Gemini 3.1 Pro (Preview)88.09
3Claude Opus 4.687.64
4Gemini 3.5 Flash87.43
5Gemini 3 Pro (Preview)87.33
6Gemini 3 Flash (Preview)86.91
7Claude Opus 4.886.37
8Claude Opus 4.5 (Non-reasoning)86.15
9GPT-5.586.06
10Claude Opus 4.586.05
11Grok 486.02
12Kimi K385.74
13Qwen 3.7 Max85.44
14Seed 2.1 Pro85.07
15Kimi K2.684.87

Interactive version: theaggregate.ai/benchmark?slug=agi-eval-community-cognition · How It Works · Data refreshed daily, snapshot 2026-09-19.