OLMo-2-1124-13B-Instruct: benchmark results

Provider: Allen AI. Released 2024-11-26. Access: Open.

Unified ELO 1428 ± 24, rank #1097 of 1639 rated models, from 145 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EuroEval English NLU - SST-569.45Sentiment classification Score (%)94.9
HREF35.6Average HREF Score (%)78.8
EuroEval English NLU - SQuAD83.61Reading comprehension Score (%)77.1
EuroEval Portuguese NLU - MultiWikiQA PT72.88Reading comprehension Score (%)73.3
Open PL LLM - PolEmo2-OUT (multiple choice, 5-shot)73.68Accuracy (%)72.9
EuroEval French NLU - Allocine94.09Sentiment classification Score (%)72.1
EuroEval English NLU67.97NLU Average Score (%)69.4
EuroEval English68.01Average Score (%)67.2
EuroEval English NLU - ScaLA EN48.67Linguistic acceptability Score (%)66.4
EuroEval Dutch NLU - SQuAD NL74.61Reading comprehension Score (%)65.4
EuroEval Italian NLU - SQuAD IT69.44Reading comprehension Score (%)65.3
EuroEval Spanish NLU - Sentiment Headlines ES43.62Sentiment classification Score (%)64.1

Interactive version: theaggregate.ai/model?slug=olmo-2-1124-13b-instruct · How It Works · Data refreshed daily, snapshot 2026-10-09.