EuroLLM-22B-Instruct-2512: benchmark results

The EU-funded EuroLLM project's 22B instruct model covering the 24 official EU languages plus 11 others, trained on EuroHPC compute (December 2025). Provider: EuroLLM. Released 2025-12-01. Access: Open.

Unified ELO 1458 ± 1, rank #1884 of 3078 rated models, from 328 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open PL LLM - PolEmo2-OUT (multiple choice, 5-shot)81.98Accuracy (%)98.4
EuroEval Croatian NLU - MMS HR46.07Sentiment classification Score (%)96.7
EuroEval German NLU - ScaLA DE65.65Linguistic acceptability Score (%)96.6
EuroEval Greek NLU - ScaLA EL57.12Linguistic acceptability Score (%)96.1
EuroEval Italian NLU - ScaLA IT59.27Linguistic acceptability Score (%)96.1
EuroEval Estonian NLU - Grammar ET51.64Linguistic acceptability Score (%)96
EuroEval Portuguese NLU - ScaLA PT55.09Linguistic acceptability Score (%)95
Open PL LLM - PPC (multiple choice, 5-shot)80.3Accuracy (%)93.3
EuroEval Spanish NLU - Sentiment Headlines ES50.66Sentiment classification Score (%)92.7
EuroEval Danish Summarization - Nordjylland News37.44Score (%)92.4
EuroEval Estonian NLU - Estonian Valence62.13Sentiment classification Score (%)92
EuroEval Polish NLU - ScaLA PL55.16Linguistic acceptability Score (%)91.4

Interactive version: theaggregate.ai/model?slug=eurollm-22b-instruct-2512 · How It Works · Data refreshed daily, snapshot 2026-09-19.