Qwen2.5-Math-72B-Instruct — benchmark results
Alibaba's 72B open math specialist (September 2024) using chain-of-thought and tool-integrated reasoning, hitting 92.9 on MATH with TIR. Provider: Alibaba. Released 2024-09-16. Access: Open.
Unified ELO 1493 ± 39, rank #827 of 1776 rated models, from 24 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open LLM Leaderboard - MATH Level 5 | 62.39 | Score | 99.9 |
| Open LLM Leaderboard - BBH | 48.97 | Score | 89.6 |
| Open LLM Leaderboard - MMLU-Pro | 42.36 | Score | 86.7 |
| Omni-MATH | 36.2 | Overall Accuracy (%) | 85.7 |
| Open LLM Leaderboard - MuSR | 16.34 | Score | 85.5 |
| Open LLM Leaderboard - GPQA | 10.85 | Score | 81 |
| U-MATH - Sequences & Series | 68.83 | Accuracy (%) | 78.8 |
| U-MATH - Integral Calculus | 38.94 | Accuracy (%) | 77.3 |
| Open FinLLM Reasoning - FinQA | 69.74 | Accuracy (%) | 76 |
| U-MATH - Multivariable Calculus | 61.8 | Accuracy (%) | 75.8 |
| U-MATH | 59.45 | Accuracy (%) | 72.7 |
| U-MATH - Precalculus | 84.38 | Accuracy (%) | 68.2 |
Interactive version: theaggregate.ai/model?slug=qwen2-5-math-72b-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.