Kimi K1.5: benchmark results

Provider: Moonshot. Access: API.

Unified ELO 1653 ± 75, rank #448 of 2656 rated models, from 13 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LLM Stats (MathVista)74.9Score (%)81.6
SuperCLUE General (May 2025) - Math Reasoning59.68Score75.6
SuperCLUE General (May 2025) - Agent50.68Score56.4
ZeroEval MATH-50096.2MATH-500 Score54.7
SuperCLUE General (May 2025) - Precise Instruction Following25.77Score53.8
SuperCLUE General (May 2025) - Overall53.72Score51.3
MathVision38.6Overall Accuracy (%)50
LLM Stats (C-Eval)88.3Score (%)47.4
SuperCLUE General (May 2025) - Science Reasoning38.61Score44.9
LLM Stats Score17.16LLM Stats Score (conservative rating)40.5
SuperCLUE General (May 2025) - Code Generation73.37Score38.5
LLM Stats (AIME 2024)77.5Score (%)36.8

Interactive version: theaggregate.ai/model?slug=kimi-k1-5 · How It Works · Data refreshed daily, snapshot 2026-09-19.