K2-V2-Instruct — benchmark results
LLM360's instruction-tuned K2-V2, a fully open 70B dense model with selectable reasoning effort. Provider: LLM360. Released 2026-01-27. Access: Open.
Unified ELO 1654 ± 27, rank #359 of 1776 rated models, from 15 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| GSM-MC | 99.17 | Accuracy (%) | 71.6 |
| RewardBench 2 Focus | 90.3 | Accuracy (%) | 69.6 |
| JudgeBench Reasoning | 93.88 | Accuracy (%) | 68.6 |
| MATH-MC Level 1 | 98.37 | Accuracy (%) | 64.7 |
| RewardBench 2 Safety | 86.67 | Accuracy (%) | 58.8 |
| RewardBench 2 Factuality | 73.95 | Accuracy (%) | 56.9 |
| RewardBench 2 Precise IF | 59.38 | Accuracy (%) | 52.9 |
| RewardBench 2 Math | 86.34 | Accuracy (%) | 51 |
| MATH-MC Level 5 | 97.36 | Accuracy (%) | 42.6 |
| MATH-MC Level 2 | 97.96 | Accuracy (%) | 41.9 |
| MATH-MC Level 4 | 98 | Accuracy (%) | 41.2 |
| MATH-MC Level 3 | 98.21 | Accuracy (%) | 39.7 |
Interactive version: theaggregate.ai/model?slug=k2-v2-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.