llama-3-gutenberg-8B: benchmark results

Provider: Meta. Access: Open.

Unified ELO 1506 ± 19, rank #1152 of 2928 rated models, from 12 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard v1 - MMLU68.57Accuracy (%) (5-shot)91.6
Open LLM Leaderboard v1 - GSM8K69.98Accuracy (%) (5-shot)91.1
Open LLM Leaderboard v1 - ARC Challenge71.33Normalized accuracy (%) (25-shot)89
Open LLM Leaderboard v1 - HellaSwag86.47Normalized accuracy (%) (10-shot)81.2
Open LLM Leaderboard v1 - TruthfulQA MC263.58MC2 (%) (0-shot)80.9
Open LLM Leaderboard v1 - WinoGrande79.16Accuracy (%) (5-shot)67.7
Open LLM Leaderboard - MMLU-Pro38.31Score67.3
Open LLM Leaderboard - GPQA30.12Score57.4
Open LLM Leaderboard - BBH49.94Score48.2
Open LLM Leaderboard - MuSR40.73Score48.1
Open LLM Leaderboard - IFEval43.72Score46.6
Open LLM Leaderboard - MATH Level 57.85Score40.4

Interactive version: theaggregate.ai/model?slug=llama-3-gutenberg-8b · How It Works · Data refreshed daily, snapshot 2026-09-23.