Kimi K3 (High): benchmark results

Provider: Moonshot. Released 2026-07-16. Access: Open.

Unified ELO 1736 ± 1, rank #56 of 3078 rated models, from 30 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AI Coding Daily (OpenCode) - Laravel Code Quality18.4Laravel Code Quality (max 20) points, LLM-judged rubric scor100
AI Coding Daily (OpenCode) - Total50.85Total points (max 60)100
Epoch AI - GPQA Diamond91.92Accuracy (%)92.9
OckBench87.5Accuracy (%)86.4
Maze-Bench48.14Mean Score84.4
OTIS Mock AIME 2024-2593.33Accuracy (%)83
Deep20Bench12.74Questions to Solve (mean, lower is better)82.4
OpenCompass Math - Competition68.9Score (%)75
Chess Puzzles (Epoch AI)25Accuracy (%)72.1
OpenCompass Knowledge - Social Science90.4Score (%)71.4
ARC-AGI-186.67Accuracy (%)70
ARC-AGI-255Accuracy (%)70

Interactive version: theaggregate.ai/model?slug=kimi-k3-high · How It Works · Data refreshed daily, snapshot 2026-09-19.