Kimi K2 (Thinking): benchmark results
Kimi K2 evaluated with thinking enabled. Provider: Moonshot. Released 2025-11-06. Access: Open.
Unified ELO 1586 ± 1, rank #488 of 1761 rated models, from 272 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CLEM Wordle | 73 | Game Clemscore (%) | 100 |
| AA AIME 2025 | 94.67 | Accuracy (%) | 97.2 |
| AA LiveCodeBench | 85.29 | Pass@1 (%) | 96.7 |
| CLEM TextMapWorld GraphReasoning | 84.28 | Game Clemscore (%) | 96.2 |
| AA Omniscience - Software Engineering (SWE) - Kotlin | 46 | Accuracy (%) | 96.1 |
| AA Omniscience - Software Engineering (SWE) - Dart | 40 | Accuracy (%) | 94.5 |
| UGI Leaderboard | 52.68 | UGI Score | 94 |
| UGI - Writing | 59.28 | Writing Score | 92.6 |
| AI for Education Pedagogy - Technology | 86.79 | Accuracy (%) | 92.4 |
| AA Omniscience - Software Engineering (SWE) - C | 63 | Accuracy (%) | 92 |
| AA Omniscience - Software Engineering (SWE) - TypeScript | 47.78 | Accuracy (%) | 92 |
| UGI - Natural Intelligence | 54.36 | NatInt Score | 91.6 |
Interactive version: theaggregate.ai/model?slug=kimi-k2-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.