AGI-Eval Community - Cognition (Chinese): leaderboard

Metric: Accuracy (%). Source: agi-eval.cn. 139 models tracked.

Top models

#ModelScore
1Claude Fable 589.64
2Gemini 3 Pro (Preview)89.11
3Gemini 3.1 Pro (Preview)88.9
4GPT-5.588
5Gemini 3.5 Flash87.85
6Qwen 3.7 Max87.09
7Claude Opus 4.687.08
8Claude Opus 4.886.87
9Kimi K386.83
10Gemini 3 Flash (Preview)86.8
11Claude Opus 4.5 (Non-reasoning)85.4
12Seed 2.1 Pro85.2
13Grok 484.86
14Kimi K2.684.24
15GPT-5.484.12

Interactive version: theaggregate.ai/benchmark?slug=agi-eval-community-cognition-chinese · How It Works · Data refreshed daily, snapshot 2026-09-19.