Kimi K2 — benchmark results

Moonshot Kimi K2 open 1T-parameter MoE model (32B active). Provider: Moonshot. Released 2025-07-11. Access: Open.

Unified ELO 1635 ± 6, rank #402 of 1776 rated models, from 448 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM Capabilities - WildBench86.19WB Score100
SciEval - Earth Sciences77.77Earth Sciences (%)100
UGI Leaderboard56.55UGI Score97.3
HELM Capabilities - Omni-MATH65.4Acc96
ProLLM - Summarization95.9Score (%)94
HELM Safety97.9Mean score (self-reported)93.4
Galileo Agent - Banking Accuracy58Accuracy (%)92.9
ProLLM - OpenBook Q&A86.3Score (%)92.4
Galileo Agent - Investment TSQ93Task Success Quality (%)90.5
Galileo Agent - Telecom TSQ91Task Success Quality (%)90.5
UGI - Natural Intelligence48.87NatInt Score90.5
Gorilla API Bench (BFCL)59.06Overall Accuracy (%)90.2

Interactive version: theaggregate.ai/model?slug=kimi-k2 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.