AGI-Eval Community - Interaction: leaderboard

Metric: Accuracy (%). Source: agi-eval.cn. 141 models tracked.

Top models

#ModelScore
1Claude Opus 4.891.2
2Claude Fable 590.94
3GPT-5.590.65
4Claude Opus 4.590.07
5Claude Opus 4.690.06
6Seed 2.0 Pro89.97
7Gemini 3.1 Pro (Preview)89.95
8GPT-5.489.74
9Kimi K389.52
10Gemini 3.5 Flash89.48
11Seed 2.1 Pro89.32
12Qwen 3.7 Max88.82
13Gemini 3 Flash (Preview)88.41
14Claude Sonnet 587.58
15Kimi K2.687.52

Interactive version: theaggregate.ai/benchmark?slug=agi-eval-community-interaction · How It Works · Data refreshed daily, snapshot 2026-09-19.