Kimi K2.5 (Thinking) — benchmark results

Kimi K2.5 evaluated with thinking enabled. Provider: Moonshot. Released 2026-01-27. Access: Open.

Unified ELO 1788 ± 11, rank #153 of 1776 rated models, from 115 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
UGI Leaderboard55.23UGI Score96.4
AA TAU-2 Bench95.91Accuracy (%)96.2
UGI - Natural Intelligence62.96NatInt Score95.8
Vals AI CorpFin v268.26Accuracy (%)95.1
UGI - Writing62.42Writing Score94.7
Kagi LLM Benchmark78.5Accuracy (%)93.6
BenchTable74.3Total Score (%)92.4
LLMEval-Logic Formalization Fixed59.4Accuracy (%)92.3
MathVision85Overall Accuracy (%)91.9
AA Omniscience - Software Engineering (SWE) - Swift72Accuracy (%)91.7
AA SciCode48.96Accuracy (%)91.1
Vals AI AIME95.62Accuracy (%)91.1

Interactive version: theaggregate.ai/model?slug=kimi-k2-5-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.