Kimi K2.6 (Thinking) — benchmark results

Moonshot Kimi K2.6 evaluated with thinking enabled. Provider: Moonshot. Released 2026-04-20. Access: Open.

Unified ELO 1791 ± 21, rank #147 of 1776 rated models, from 62 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Qwen3.7 Launch - Humanity's Last Exam (with tools)54Score (%)100
Qwen3.7 Launch - MMLU-Redux95.3Score (%)100
UGI - Natural Intelligence65.25NatInt Score96.7
Wolfram LLM Benchmarking Project67.3Correct Functionality (%)96.2
UGI - Writing64.4Writing Score95.8
UGI Leaderboard51.08UGI Score91.7
Qwen3.7 Launch - IFEval94.5Score (%)90
MathArena - ArXiv Math Jan 202671.74Accuracy (%)84.6
Qwen3.7 Launch - PolyMATH82.7Score (%)80
Qwen3.7 Launch - SWE-Pro59.5Resolved (%)80
LLM2014 Logic 2026-0457.33Median Score77.5
Qwen3.7 Launch - SciCode52.2Score (%)75

Interactive version: theaggregate.ai/model?slug=kimi-k2-6-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.