K-12EduBench — leaderboard
K-12 education benchmark for subject knowledge, problem solving, and educational-goal cognition across school-level tasks.
Metric: Lambda_js Avg (self-reported). Source: benchmarklist.com. Status: saturation imminent. 23 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek V3 | 79.67 |
| 2 | Yi Lightning | 69.87 |
| 3 | Gemini 1.5 Pro | 67.88 |
| 4 | Grok 2 | 64.1 |
| 5 | Gemini 2.0 Flash | 64.05 |
| 6 | Grok 3 | 63.91 |
| 7 | Claude 3.7 Sonnet | 61.2 |
| 8 | GPT-4 Turbo | 55.94 |
| 9 | O1 Mini | 54.46 |
| 10 | Llama 3.1 70B | 49.65 |
| 11 | Claude 3.5 Haiku | 44.98 |
| 12 | Llama 3.1 8B | 21.91 |
Interactive version: theaggregate.ai/benchmark?slug=k-12edubench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.