Kimi K2.6 (Thinking) — benchmark results
Moonshot Kimi K2.6 evaluated with thinking enabled. Provider: Moonshot. Released 2026-04-20. Access: Open.
Unified ELO 1791 ± 21, rank #147 of 1776 rated models, from 62 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Qwen3.7 Launch - Humanity's Last Exam (with tools) | 54 | Score (%) | 100 |
| Qwen3.7 Launch - MMLU-Redux | 95.3 | Score (%) | 100 |
| UGI - Natural Intelligence | 65.25 | NatInt Score | 96.7 |
| Wolfram LLM Benchmarking Project | 67.3 | Correct Functionality (%) | 96.2 |
| UGI - Writing | 64.4 | Writing Score | 95.8 |
| UGI Leaderboard | 51.08 | UGI Score | 91.7 |
| Qwen3.7 Launch - IFEval | 94.5 | Score (%) | 90 |
| MathArena - ArXiv Math Jan 2026 | 71.74 | Accuracy (%) | 84.6 |
| Qwen3.7 Launch - PolyMATH | 82.7 | Score (%) | 80 |
| Qwen3.7 Launch - SWE-Pro | 59.5 | Resolved (%) | 80 |
| LLM2014 Logic 2026-04 | 57.33 | Median Score | 77.5 |
| Qwen3.7 Launch - SciCode | 52.2 | Score (%) | 75 |
Interactive version: theaggregate.ai/model?slug=kimi-k2-6-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.