Kimi K3 (High): benchmark results
Provider: Moonshot. Released 2026-07-16. Access: Open.
Unified ELO 1736 ± 1, rank #56 of 3078 rated models, from 30 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AI Coding Daily (OpenCode) - Laravel Code Quality | 18.4 | Laravel Code Quality (max 20) points, LLM-judged rubric scor | 100 |
| AI Coding Daily (OpenCode) - Total | 50.85 | Total points (max 60) | 100 |
| Epoch AI - GPQA Diamond | 91.92 | Accuracy (%) | 92.9 |
| OckBench | 87.5 | Accuracy (%) | 86.4 |
| Maze-Bench | 48.14 | Mean Score | 84.4 |
| OTIS Mock AIME 2024-25 | 93.33 | Accuracy (%) | 83 |
| Deep20Bench | 12.74 | Questions to Solve (mean, lower is better) | 82.4 |
| OpenCompass Math - Competition | 68.9 | Score (%) | 75 |
| Chess Puzzles (Epoch AI) | 25 | Accuracy (%) | 72.1 |
| OpenCompass Knowledge - Social Science | 90.4 | Score (%) | 71.4 |
| ARC-AGI-1 | 86.67 | Accuracy (%) | 70 |
| ARC-AGI-2 | 55 | Accuracy (%) | 70 |
Interactive version: theaggregate.ai/model?slug=kimi-k3-high · How It Works · Data refreshed daily, snapshot 2026-09-19.