Llama 3.1 8B Cobalt — benchmark results

Valiant Labs' math-instruct fine-tune of Llama 3.1 8B Instruct, trained on synthetic math data generated with Llama 3.1 405B. Provider: Meta. Released 2024-08-16. Access: Open.

Unified ELO 1467 ± 34, rank #925 of 1776 rated models, from 7 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard - IFEval71.68Score85.2
Open Korean LLM Leaderboard368.64Average Score (%)73.4
Open LLM Leaderboard - MATH Level 515.33Score62.5
Open LLM Leaderboard - GPQA7.16Score60.3
Open LLM Leaderboard - MMLU-Pro29.59Score57.6
Open LLM Leaderboard - MuSR9.83Score48.4
Open LLM Leaderboard - BBH27.42Score43.7

Interactive version: theaggregate.ai/model?slug=llama-3-1-8b-cobalt · How the rankings work · Data refreshed daily, snapshot 2026-07-22.