LLMEval-Fair - Education: leaderboard

Metric: Discipline Score (0-10). Source: github.com. 59 models tracked.

Top models

#ModelScore
1DeepSeek R19.27
2QwQ-32B9.23
3Gemini 2.5 Pro (Preview)9.2
4Qwen 3 235B A22B9.03
5Claude Sonnet 4.58.93
6DeepSeek V38.93
7GPT-58.9
8Kimi K28.8
9Claude Sonnet 4.5 (Thinking)8.8
10GLM-4.68.7
11Gemini 2.5 Flash (Thinking)8.7
12O1 (2024-12-17)8.67
13Claude Sonnet 4 (Thinking)8.63
14Qwen 3 32B8.57
15DeepSeek V3.28.53

Interactive version: theaggregate.ai/benchmark?slug=llmeval-fair-education · How It Works · Data refreshed daily, snapshot 2026-09-19.