Qwen2.5-Math-7B: benchmark results

Provider: Alibaba. Released 2024-09-16. Access: Open.

Unified ELO 1396 ± 1, rank #1152 of 1392 rated models, from 47 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
PhysicsFinals100Score (self-reported)100
Open LLM Leaderboard - MATH Level 530.51Score82.1
EuroEval Spanish NLU - Sentiment Headlines ES40.68Sentiment classification Score (%)54.1
Open LLM Leaderboard - GPQA5.82Score48.7
EuroEval Spanish NLU - ScaLA ES9.62Linguistic acceptability Score (%)45.3
EuroEval Spanish NLU39.78NLU Average Score (%)42.4
EuroEval Portuguese NLU - ScaLA PT7.68Linguistic acceptability Score (%)40.7
EuroEval Spanish34.09Average Score (%)39.7
EuroEval Portuguese37.62Average Score (%)39.2
EuroEval Portuguese NLU - SST-2 PT72.07Sentiment classification Score (%)38.2
EuroEval Norwegian Common Sense Reasoning29.52Common Sense Reasoning Average Score (%)38.1
EuroEval Portuguese Common Sense Reasoning19.35Common Sense Reasoning Average Score (%)38.1

Interactive version: theaggregate.ai/model?slug=qwen2-5-math-7b · How It Works · Data refreshed daily, snapshot 2026-09-05.