Kimi K2.6 (Thinking): benchmark results
Moonshot Kimi K2.6 evaluated with thinking enabled. Provider: Moonshot. Released 2026-04-20. Access: Open.
Unified ELO 1645 ± 1, rank #258 of 1761 rated models, from 106 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLM Stats (MMLU-Redux) | 95.3 | Score (%) | 100 |
| Qwen3.7 Launch - Humanity's Last Exam (with tools) | 54 | Score (%) | 100 |
| SuperCLUE-Writing - Content Creativity | 89.39 | Score | 100 |
| UGI - Natural Intelligence | 65.25 | NatInt Score | 96.1 |
| UGI - Writing | 64.4 | Writing Score | 95.1 |
| Wolfram LLM Benchmarking Project | 67.3 | Correct Functionality (%) | 94.9 |
| VitaBench 2.0 | 48.1 | Avg@4 Full Context (self-reported) | 94.1 |
| LLM Stats (PolyMATH) | 82.7 | Score (%) | 92.3 |
| SuperCLUE-Writing - Content Style | 90.96 | Score | 92.3 |
| UGI Leaderboard | 51.08 | UGI Score | 91.6 |
| Qwen3.7 Launch - IFEval | 94.5 | Score (%) | 90 |
| AtmosCoder-Bench | 93.8 | Accuracy (%, mean of 3 runs) | 86.7 |
Interactive version: theaggregate.ai/model?slug=kimi-k2-6-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.