Kimi K2.5: benchmark results
Moonshot Kimi K2.5 open 1T-parameter MoE model. Provider: Moonshot. Released 2026-01-27. Access: Open.
Unified ELO 1658 ± 1, rank #112 of 1392 rated models, from 571 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AGC-Bench - Problem Solving | 1.02 | JRT z-score | 100 |
| AGC-Bench - rpgbench | 1.74 | Dataset z-score | 100 |
| AGC-Bench - tinyfabulist | 2.03 | Dataset z-score | 100 |
| Diagram-MMU - Code Parsing | 56.97 | All Score (%) | 100 |
| EHR-Complex | 62.3 | Avg. (self-reported) | 100 |
| FINAL Bench Metacognitive | 78.54 | Metacognitive Score (self-reported) | 100 |
| LLM Stats (InfoVQAtest) | 92.6 | Score (%) | 100 |
| LLM Stats (MathVista-Mini) | 90.1 | Score (%) | 100 |
| LLM Stats (OCRBench) | 92.3 | Score (%) | 100 |
| LLM Stats (Seal-0) | 57.4 | Score (%) | 100 |
| MATH-MC Level 3 | 99.73 | Accuracy (%) | 100 |
| MERA Code - UnitTests | 34.21 | CodeBLEU (%) | 100 |
Interactive version: theaggregate.ai/model?slug=kimi-k2-5 · How It Works · Data refreshed daily, snapshot 2026-09-05.