Llama-3-Instruct-8B-SimPO-ExPO — benchmark results

chujiezheng's ExPO weight extrapolation of Princeton NLP's SimPO-tuned Llama 3 8B Instruct, nudging alignment beyond the original checkpoint. Provider: Meta. Released 2024-05-26. Access: Open.

Unified ELO 1452 ± 33, rank #990 of 1776 rated models, from 8 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AlpacaEval 2.045.78LC Win Rate (%)84.7
Open LLM Leaderboard - IFEval64.34Score76.8
Open LLM Leaderboard - MMLU-Pro26.68Score48.9
Open LLM Leaderboard - MuSR9.5Score46.3
Open LLM Leaderboard - GPQA4.92Score41.7
WildBench35.02WB Score Task-Macro40.3
Open LLM Leaderboard - BBH25.87Score40
Open LLM Leaderboard - MATH Level 57.02Score37.5

Interactive version: theaggregate.ai/model?slug=llama-3-instruct-8b-simpo-expo · How the rankings work · Data refreshed daily, snapshot 2026-07-22.