K2-Chat — benchmark results

LLM360's chat fine-tune of its fully open 65B K2 model, a reproducible LLM claimed to beat Llama 2 70B Chat using 35% less compute (2024). Provider: LLM360. Released 2024-05-22. Access: Open.

Unified ELO 1472 ± 21, rank #904 of 1776 rated models, from 7 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard - MuSR16.82Score86.9
Open LLM Leaderboard - BBH33.79Score66.4
Open LLM Leaderboard - GPQA7.49Score63.1
Open LLM Leaderboard - IFEval51.52Score61.4
Open LLM Leaderboard - MATH Level 510.35Score48.5
Open LLM Leaderboard - MMLU-Pro26.34Score48.1
SuperGPQA22.47Accuracy (%)17.6

Interactive version: theaggregate.ai/model?slug=k2-chat · How the rankings work · Data refreshed daily, snapshot 2026-07-22.