AGI-Eval Community - Subject Reasoning (English): leaderboard

Metric: Accuracy (%). Source: agi-eval.cn. 139 models tracked.

Top models

#ModelScore
1Claude Opus 4.888.79
2GPT-5.587.52
3Claude Fable 587.32
4Claude Opus 4.687.21
5Gemini 3.5 Flash87.16
6Seed 2.0 Pro87.11
7Qwen 3.7 Max87.05
8Kimi K2.5 (Thinking)86.9
9Kimi K386.89
10Kimi K2.686.79
11Gemini 3 Pro (Preview)86.63
12Gemini 3.1 Pro (Preview)86.53
13Seed 2.1 Pro86.37
14GPT-5.186.16
15LongCat 2.085.95

Interactive version: theaggregate.ai/benchmark?slug=agi-eval-community-subject-reasoning-english · How It Works · Data refreshed daily, snapshot 2026-09-19.