K2-V2-Instruct — benchmark results

LLM360's instruction-tuned K2-V2, a fully open 70B dense model with selectable reasoning effort. Provider: LLM360. Released 2026-01-27. Access: Open.

Unified ELO 1654 ± 27, rank #359 of 1776 rated models, from 15 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
GSM-MC99.17Accuracy (%)71.6
RewardBench 2 Focus90.3Accuracy (%)69.6
JudgeBench Reasoning93.88Accuracy (%)68.6
MATH-MC Level 198.37Accuracy (%)64.7
RewardBench 2 Safety86.67Accuracy (%)58.8
RewardBench 2 Factuality73.95Accuracy (%)56.9
RewardBench 2 Precise IF59.38Accuracy (%)52.9
RewardBench 2 Math86.34Accuracy (%)51
MATH-MC Level 597.36Accuracy (%)42.6
MATH-MC Level 297.96Accuracy (%)41.9
MATH-MC Level 498Accuracy (%)41.2
MATH-MC Level 398.21Accuracy (%)39.7

Interactive version: theaggregate.ai/model?slug=k2-v2-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.