Kimi K2 (0711) — benchmark results

July 11, 2025 Kimi K2 snapshot, tracked when sources report the dated API model. Provider: Moonshot. Released 2025-07-11. Access: Open.

Unified ELO 1514 ± 29, rank #751 of 1776 rated models, from 27 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SuperGPQA58.08Accuracy (%)79.4
Chatbot Arena (Text)1417Elo72.5
LLM2014 Code 2025-09 - C++4.65Score57.9
LMGame-Bench Tetris17Score54.2
Multi-Docker-Eval34.23Resolved (%)53.3
LLM2014 Code 2025-09 - Java5.1Score47.4
MCPMark19.09Pass@1 (%)39.5
ConStory-Bench1.33CED errors per 10K words (lower is better)37.5
Epoch AI - ECI140.45ECI Score36.4
LiveSecBench35.58Overall Score (%)33.3
Kagi LLM Benchmark45Accuracy (%)32.9
Design Arena (UI Components)1073Elo22.8

Interactive version: theaggregate.ai/model?slug=kimi-k2-0711 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.