Kimi K2.6 (Thinking): benchmark results

Moonshot Kimi K2.6 evaluated with thinking enabled. Provider: Moonshot. Released 2026-04-20. Access: Open.

Unified ELO 1645 ± 1, rank #258 of 1761 rated models, from 106 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LLM Stats (MMLU-Redux)95.3Score (%)100
Qwen3.7 Launch - Humanity's Last Exam (with tools)54Score (%)100
SuperCLUE-Writing - Content Creativity89.39Score100
UGI - Natural Intelligence65.25NatInt Score96.1
UGI - Writing64.4Writing Score95.1
Wolfram LLM Benchmarking Project67.3Correct Functionality (%)94.9
VitaBench 2.048.1Avg@4 Full Context (self-reported)94.1
LLM Stats (PolyMATH)82.7Score (%)92.3
SuperCLUE-Writing - Content Style90.96Score92.3
UGI Leaderboard51.08UGI Score91.6
Qwen3.7 Launch - IFEval94.5Score (%)90
AtmosCoder-Bench93.8Accuracy (%, mean of 3 runs)86.7

Interactive version: theaggregate.ai/model?slug=kimi-k2-6-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.