Kimi K2 Instruct (0905): benchmark results
Provider: Moonshot. Released 2025-09-05. Access: Open.
Unified ELO 1744 ± 22, rank #252 of 2656 rated models, from 35 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| TuRTLe - Icarus Syntax | 94.22 | Average Score (%) | 95.5 |
| TuRTLe - Verilator Syntax | 95.09 | Average Score (%) | 93.2 |
| BALSAM - Text Classification | 40.31 | Overall score (0-100, LLM-judged generation and multiple cho | 85.7 |
| Wolfram LLM Benchmarking Project | 57.2 | Correct Functionality (%) | 84.4 |
| TuRTLe - Verilator Performance | 67.03 | Average Score (%) | 84.1 |
| TuRTLe - Verilator Synthesis | 68.68 | Average Score (%) | 84.1 |
| BALSAM - Creative Writing | 51.19 | Overall score (0-100, LLM-judged generation and multiple cho | 82.1 |
| TuRTLe - Icarus Performance | 67.47 | Average Score (%) | 81.8 |
| TuRTLe - Icarus Synthesis | 69.25 | Average Score (%) | 81.8 |
| TuRTLe Code Completion (Icarus Verilog) | 71.77 | Aggregated Score (self-reported) | 81.4 |
| TuRTLe Code Completion (Verilator) | 71.79 | Aggregated Score (self-reported) | 81.4 |
| TuRTLe Spec-to-RTL (Icarus Verilog) | 68.72 | Aggregated Score (self-reported) | 81.4 |
Interactive version: theaggregate.ai/model?slug=kimi-k2-instruct-0905 · How It Works · Data refreshed daily, snapshot 2026-09-19.