Kimi K3 (Thinking): benchmark results
Provider: Moonshot. Released 2026-07-16. Access: Open.
Unified ELO 1732 ± 1, rank #64 of 3078 rated models, from 86 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Nejumi 4 - jaster (2-shot) - JaMP | 85 | Exact match (%) | 99.3 |
| Nejumi 4 - MT-Bench (Japanese) - Writing | 99.5 | Judge rating (1-10, x10) | 98.9 |
| Nejumi 4 - jaster (0-shot) - MMLU-ProX (Japanese) | 91 | Exact match (%) | 97.8 |
| MathArena - Kangaroo 2025 Levels 7-8 | 96.67 | Accuracy (%) | 97.6 |
| Nejumi 4 - jaster (2-shot) - MMLU-ProX (Japanese) | 91 | Exact match (%) | 97.1 |
| Wolfram LLM Benchmarking Project | 68.8 | Correct Functionality (%) | 97.1 |
| Nejumi 4 - jaster (0-shot) - JMMLU | 96 | Exact match (%) | 96.7 |
| MathArena - APEX Shortlist 2025 | 92.02 | Accuracy (%) | 95.9 |
| Nejumi 4 - HalluLens | 100 | Hallucination resistance (%) | 95.6 |
| Nejumi 4 - HLE (Japanese) - Accuracy | 44.33 | Accuracy (%) | 94.1 |
| ProfBench | 57.9 | Overall Rubric Score (%) | 94.1 |
| Nejumi 4 - Toxicity - Fairness | 99.44 | Criteria met (%) | 93.8 |
Interactive version: theaggregate.ai/model?slug=kimi-k3-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.