K2 Think V2 — benchmark results
MBZUAI's open sovereign 70B reasoning model, RL-trained on the K2-V2 base. Provider: LLM360. Released 2026-01-08. Access: Open.
Unified ELO 1575 ± 23, rank #548 of 1776 rated models, from 37 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA IFBench | 62.79 | Accuracy (%) | 73.7 |
| AA-LCR | 52.7 | Score (self-reported) | 67.1 |
| AA Long Context Reasoning | 52.67 | Accuracy (%) | 61 |
| AA Omniscience | -33.92 | Score | 60.5 |
| AA Humanity's Last Exam | 9.45 | Accuracy (%) | 59.6 |
| AA GPQA Diamond | 71.31 | Accuracy (%) | 55.1 |
| AA Omniscience - Software Engineering (SWE) - Kotlin | 20 | Accuracy (%) | 54.9 |
| Artificial Analysis Intelligence Index | 17.26 | Intelligence Index | 53.7 |
| AA SciCode | 32.99 | Accuracy (%) | 48.5 |
| AA Omniscience - Software Engineering (SWE) - Julia | 12 | Accuracy (%) | 47.9 |
| AA Omniscience - Software Engineering (SWE) - TypeScript | 20 | Accuracy (%) | 46 |
| AA Omniscience - Software Engineering (SWE) - JavaScript | 28.18 | Accuracy (%) | 45.4 |
Interactive version: theaggregate.ai/model?slug=k2-think-v2 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.