Kimi K2.5 (High): benchmark results

Provider: Moonshot. Released 2026-01-27. Access: Open.

Unified ELO 1683 ± 1, rank #194 of 3078 rated models, from 12 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SWE-bench Verified70.8Resolved (%)76.1
FutureEval9.41Unified Forecasting Score73.9
NonoBench36.7Overall Accuracy (%)69.1
CLBench Life - Behavioral Records & Activity Trails12.8Solving Rate (%)58.9
CLBench Life13.2Solving Rate (%)57.1
CLBench Life - Communication & Social Interactions15.3Solving Rate (%)57.1
CLBench Life - Fragmented Information & Revisions11.4Solving Rate (%)53.6
SEC-bench Pro - V81.94V8 instances solved, headline mode (% of 103 instances; time30
SlopCodeBench9.69Isolated Solved (%)25
SEC-bench Pro - Firefox2.88Firefox instances solved, headline mode (% of 104 instances;20
SEC-bench Pro - Linux2.19Linux instances solved, headline mode (% of 137 instances; t20
SEC-bench Pro - Overall2.33Instances solved, headline mode (% of 344 instances across V20

Interactive version: theaggregate.ai/model?slug=kimi-k2-5-high · How It Works · Data refreshed daily, snapshot 2026-09-19.