AGI-Eval Community - Subject Reasoning (Chinese): leaderboard

Metric: Accuracy (%). Source: agi-eval.cn. 139 models tracked.

Top models

#ModelScore
1Seed 2.1 Pro94.54
2GPT-5.594.31
3Seed 2.0 Pro94.07
4Claude Fable 593.6
5Gemini 3.1 Pro (Preview)93.36
6Kimi K393.24
7Gemini 3 Pro (Preview)92.88
8Gemini 3.5 Flash92.88
9Qwen 3.7 Max92.88
10DeepSeek V4 Flash (Max)92.88
11Claude Opus 4.892.76
12Kimi K2.692.53
13DeepSeek V4 Pro (Max)92.53
14Claude Opus 4.692.41
15Seed 1.892.29

Interactive version: theaggregate.ai/benchmark?slug=agi-eval-community-subject-reasoning-chinese · How It Works · Data refreshed daily, snapshot 2026-09-19.