Kimi K2.5 — benchmark results
Moonshot Kimi K2.5 open 1T-parameter MoE model. Provider: Moonshot. Released 2026-01-27. Access: Open.
Unified ELO 1745 ± 11, rank #205 of 1776 rated models, from 374 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AGC-Bench - Problem Solving | 1.02 | JRT z-score | 100 |
| AGC-Bench - rpgbench | 1.74 | Dataset z-score | 100 |
| AGC-Bench - tinyfabulist | 2.03 | Dataset z-score | 100 |
| AI for Education Pedagogy - Technology | 89.62 | Accuracy (%) | 100 |
| LLM Stats (InfoVQAtest) | 92.6 | Score (%) | 100 |
| LLM Stats (MathVista-Mini) | 90.1 | Score (%) | 100 |
| LLM Stats (OCRBench) | 92.3 | Score (%) | 100 |
| LLM Stats (Seal-0) | 57.4 | Score (%) | 100 |
| MATH-MC Level 3 | 99.73 | Accuracy (%) | 100 |
| Multi-Docker-Eval | 41.82 | Resolved (%) | 100 |
| OpenEvals - Humanity's Last Exam | 50.2 | Accuracy (%) | 100 |
| RP-Leaderboard | 88.9 | RP Score (%) | 100 |
Interactive version: theaggregate.ai/model?slug=kimi-k2-5 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.