Kimi K2: benchmark results
Moonshot Kimi K2 open 1T-parameter MoE model (32B active). Provider: Moonshot. Released 2025-07-11. Access: Open.
Unified ELO 1633 ± 1, rank #171 of 1392 rated models, from 485 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Capabilities - WildBench | 86.19 | WB Score | 100 |
| MLB | 77.29 | Average (self-reported) | 100 |
| SciEval - Earth Sciences | 77.77 | Earth Sciences (%) | 100 |
| SuiChat-CN | 76.3 | Standard F1 (self-reported) | 100 |
| UGI Leaderboard | 56.55 | UGI Score | 97.4 |
| HELM Capabilities - Omni-MATH | 65.4 | Acc | 96 |
| WritingBench | 81.26 | Score (self-reported) | 94.4 |
| ProLLM - Summarization | 95.9 | Score (%) | 94 |
| HELM Safety | 97.9 | Mean score (self-reported) | 93 |
| Galileo Agent - Banking Accuracy | 58 | Accuracy (%) | 92.9 |
| ProLLM - OpenBook Q&A | 86.3 | Score (%) | 92.4 |
| AA Omniscience - Software Engineering (SWE) - Swift | 56 | Accuracy (%) | 91.5 |
Interactive version: theaggregate.ai/model?slug=kimi-k2 · How It Works · Data refreshed daily, snapshot 2026-09-05.