Llama 3.1 8B IT — benchmark results

Meta's 8B instruction-tuned Llama 3.1 model (July 2024), multilingual with 128K context and fine-tuned for tool calling. Provider: Meta. Released 2024-07-23. Access: Open.

Unified ELO 1400 ± 34, rank #1228 of 1776 rated models, from 9 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BALROG Crafter (LLM)25.5Progress (%)33.3
BALROG MiniHack (LLM)5Progress (%)31.8
BALROG BabaIsAI (LLM)18.3Progress (%)30.3
Fin-Bias92.4Average Herding Score (with rating) (self-reported)27.8
BALROG TextWorld (LLM)6.1Progress (%)24.2
BALROG BabyAI (LLM)36Progress (%)21.2
BALROG NetHack (LLM)0Progress (%)16.7
TriBench-Ko38.6Macro F1 (self-reported)8.3
LEXam24.04Multiple-Choice Accuracy (%)3.3

Interactive version: theaggregate.ai/model?slug=llama-3-1-8b-it · How the rankings work · Data refreshed daily, snapshot 2026-07-22.