Kimi K2 (0905) (Thinking) — benchmark results
Provider: Moonshot. Released 2025-07-11. Access: Open.
Unified ELO 1725 ± 35, rank #273 of 1806 rated models, from 16 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| OJBench | 48.7 | Score (self-reported) | 97.7 |
| MathArena - HMMT Feb 2025 | 97.5 | Accuracy (%) | 91.7 |
| LLM Stats (HMMT 2025) | 97.5 | Score (%) | 90.6 |
| LLM Stats (HealthBench) | 58 | Score (%) | 87.5 |
| LLM Stats (OJBench) | 48.7 | Score (%) | 87.5 |
| LLM Stats (MMLU-Redux) | 94.4 | Score (%) | 86.3 |
| LLM Stats Score | 36.78 | LLM Stats Score (conservative rating) | 82.7 |
| LLM Stats (Seal-0) | 56.3 | Score (%) | 80 |
| ZeroEval GPQA Diamond | 84.5 | GPQA Diamond Score | 74.9 |
| WritingBench | 73.8 | Score (self-reported) | 47.7 |
| LLM Stats (BrowseComp-zh) | 62.3 | Score (%) | 41.7 |
| LLM Stats (Multi-SWE-Bench) | 41.9 | Score (%) | 40 |
Interactive version: theaggregate.ai/model?slug=kimi-k2-0905-thinking · How It Works · Data refreshed daily, snapshot 2026-08-07.