Kimi K2.5 (High): benchmark results
Provider: Moonshot. Released 2026-01-27. Access: Open.
Unified ELO 1683 ± 1, rank #194 of 3078 rated models, from 12 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SWE-bench Verified | 70.8 | Resolved (%) | 76.1 |
| FutureEval | 9.41 | Unified Forecasting Score | 73.9 |
| NonoBench | 36.7 | Overall Accuracy (%) | 69.1 |
| CLBench Life - Behavioral Records & Activity Trails | 12.8 | Solving Rate (%) | 58.9 |
| CLBench Life | 13.2 | Solving Rate (%) | 57.1 |
| CLBench Life - Communication & Social Interactions | 15.3 | Solving Rate (%) | 57.1 |
| CLBench Life - Fragmented Information & Revisions | 11.4 | Solving Rate (%) | 53.6 |
| SEC-bench Pro - V8 | 1.94 | V8 instances solved, headline mode (% of 103 instances; time | 30 |
| SlopCodeBench | 9.69 | Isolated Solved (%) | 25 |
| SEC-bench Pro - Firefox | 2.88 | Firefox instances solved, headline mode (% of 104 instances; | 20 |
| SEC-bench Pro - Linux | 2.19 | Linux instances solved, headline mode (% of 137 instances; t | 20 |
| SEC-bench Pro - Overall | 2.33 | Instances solved, headline mode (% of 344 instances across V | 20 |
Interactive version: theaggregate.ai/model?slug=kimi-k2-5-high · How It Works · Data refreshed daily, snapshot 2026-09-19.