K2 Think V2: benchmark results
MBZUAI's open sovereign 70B reasoning model, RL-trained on the K2-V2 base. Provider: LLM360. Released 2026-01-27. Access: Open.
Unified ELO 1531 ± 1, rank #756 of 1919 rated models, from 35 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA IFBench | 62.79 | Accuracy (%) | 73.7 |
| AA Omniscience - Software Engineering (SWE) - JavaScript | 33.64 | Accuracy (%) | 69.1 |
| AA Omniscience - Software Engineering (SWE) - Kotlin | 20 | Accuracy (%) | 63 |
| AA Omniscience - Software Engineering (SWE) - HTML | 32 | Accuracy (%) | 60.4 |
| AA Omniscience - Software Engineering (SWE) - Rust | 50 | Accuracy (%) | 59 |
| AA Omniscience - Software Engineering (SWE) - Julia | 12 | Accuracy (%) | 58 |
| AA Omniscience - Software Engineering (SWE) - TypeScript | 21.11 | Accuracy (%) | 57.5 |
| BenchmarkList ECI | 112.49 | Capability Index (ECI) | 54.1 |
| AA Humanity's Last Exam | 10.1 | Accuracy (%) | 53.8 |
| AA Omniscience - Software Engineering (SWE) - R | 10 | Accuracy (%) | 53.4 |
| AA Long Context Reasoning | 57 | Accuracy (%) | 52.8 |
| Artificial Analysis Intelligence Index | 11.5 | Intelligence Index | 50.7 |
Interactive version: theaggregate.ai/model?slug=k2-think-v2 · How It Works · Data refreshed daily, snapshot 2026-09-08.