Qwen2.5-Math-7B — benchmark results

Provider: Alibaba. Released 2024-09-16. Access: Open.

Unified ELO 1314 ± 38, rank #1504 of 1776 rated models, from 47 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open LLM Leaderboard - MATH Level 530.51Score82.1
EuroEval Spanish NLU - Sentiment Headlines ES40.68Sentiment classification Score (%)54.1
Open LLM Leaderboard - GPQA5.82Score48.6
EuroEval Spanish NLU - ScaLA ES9.62Linguistic acceptability Score (%)45.3
EuroEval Spanish NLU39.78NLU Average Score (%)42.4
EuroEval Portuguese NLU - ScaLA PT7.68Linguistic acceptability Score (%)40.7
EuroEval Spanish34.09Average Score (%)39.7
EuroEval Portuguese37.62Average Score (%)39.2
EuroEval Portuguese NLU - SST-2 PT72.07Sentiment classification Score (%)38.2
EuroEval Norwegian Common Sense Reasoning29.52Common Sense Reasoning Average Score (%)38.1
EuroEval Portuguese Common Sense Reasoning19.35Common Sense Reasoning Average Score (%)38.1
EuroEval Spanish NLU - CoNLL ES54.68Named entity recognition Score (%)37.6

Interactive version: theaggregate.ai/model?slug=qwen2-5-math-7b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.