Kimi K2 (Thinking) — benchmark results
Kimi K2 evaluated with thinking enabled. Provider: Moonshot. Released 2025-11-06. Access: Open.
Unified ELO 1721 ± 11, rank #238 of 1776 rated models, from 236 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CLEM Wordle | 73 | Game Clemscore (%) | 100 |
| AA LiveCodeBench | 85.29 | Pass@1 (%) | 96.2 |
| CLEM TextMapWorld GraphReasoning | 84.28 | Game Clemscore (%) | 96.2 |
| AA AIME 2025 | 94.67 | Accuracy (%) | 96.1 |
| UGI Leaderboard | 52.68 | UGI Score | 94.1 |
| UGI - Writing | 59.28 | Writing Score | 93.3 |
| AI for Education Pedagogy - Technology | 86.79 | Accuracy (%) | 93 |
| UGI - Natural Intelligence | 54.36 | NatInt Score | 92.3 |
| FormationEval | 97.2 | Accuracy (%) | 91.5 |
| Wolfram LLM Benchmarking Project | 61.6 | Correct Functionality (%) | 91.3 |
| MATH-MC Level 3 | 99.28 | Accuracy (%) | 91.2 |
| MATH-MC Level 5 | 99.69 | Accuracy (%) | 90.4 |
Interactive version: theaggregate.ai/model?slug=kimi-k2-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.