Kimi K2 — benchmark results
Moonshot Kimi K2 open 1T-parameter MoE model (32B active). Provider: Moonshot. Released 2025-07-11. Access: Open.
Unified ELO 1635 ± 6, rank #402 of 1776 rated models, from 448 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Capabilities - WildBench | 86.19 | WB Score | 100 |
| SciEval - Earth Sciences | 77.77 | Earth Sciences (%) | 100 |
| UGI Leaderboard | 56.55 | UGI Score | 97.3 |
| HELM Capabilities - Omni-MATH | 65.4 | Acc | 96 |
| ProLLM - Summarization | 95.9 | Score (%) | 94 |
| HELM Safety | 97.9 | Mean score (self-reported) | 93.4 |
| Galileo Agent - Banking Accuracy | 58 | Accuracy (%) | 92.9 |
| ProLLM - OpenBook Q&A | 86.3 | Score (%) | 92.4 |
| Galileo Agent - Investment TSQ | 93 | Task Success Quality (%) | 90.5 |
| Galileo Agent - Telecom TSQ | 91 | Task Success Quality (%) | 90.5 |
| UGI - Natural Intelligence | 48.87 | NatInt Score | 90.5 |
| Gorilla API Bench (BFCL) | 59.06 | Overall Accuracy (%) | 90.2 |
Interactive version: theaggregate.ai/model?slug=kimi-k2 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.