AGI-Eval Community - Instruction Following: leaderboard

Metric: Accuracy (%). Source: agi-eval.cn. 141 models tracked.

Top models

#ModelScore
1GPT-5.571.29
2GPT-5.469.47
3GLM-5.269.23
4Qwen 3.7 Max68.99
5Gemini 3.5 Flash67.09
6Seed 2.1 Pro67.09
7Kimi K2.665.74
8O3 Pro64.63
9Kimi K364.23
10GPT-5.164.08
11DeepSeek V4 Pro (Max)64.08
12Gemini 3.1 Pro (Preview)64
13Claude Opus 4.863.52
14Seed 2.0 Pro63.52
15GPT-5 (High)63.28

Interactive version: theaggregate.ai/benchmark?slug=agi-eval-community-instruction-following · How It Works · Data refreshed daily, snapshot 2026-09-19.