Qwen 2.5 32B Instruct — benchmark results
Alibaba Qwen 2.5 32B instruction-tuned checkpoint. Provider: Alibaba. Released 2024-09-19. Access: Open.
Unified ELO 1559 ± 11, rank #594 of 1776 rated models, from 363 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| PlantMarkerBench | 75.4 | Valid F1 (self-reported) | 100 |
| SeaEval - Fundamental NLP Tasks - MNLI (Zero-Shot) | 87.15 | Accuracy (%) | 100 |
| SeaEval - Fundamental NLP Tasks - RTE (Zero-Shot) | 90.97 | Accuracy (%) | 100 |
| Open LLM Leaderboard - MATH Level 5 | 62.54 | Score | 99.9 |
| Open LLM Leaderboard - IFEval | 83.46 | Score | 99.4 |
| Open LLM Leaderboard - MMLU-Pro | 51.85 | Score | 98.9 |
| Open LLM Leaderboard - BBH | 56.49 | Score | 98.5 |
| OpenEval - BBQ | 95.3 | Exact Match (%) | 98.3 |
| Open Arabic LLM - Arabic MMLU HT Moral Scenarios | 58.1 | Accuracy (%) | 98.1 |
| MT-Bench PL - Reasoning | 9.1 | Judge Score (0-10) | 98 |
| SeaEval - Fundamental NLP Tasks - QQP (Zero-Shot) | 83.15 | Accuracy (%) | 97.8 |
| MT-Bench PL - Extraction | 9.9 | Judge Score (0-10) | 96.9 |
Interactive version: theaggregate.ai/model?slug=qwen-2-5-32b-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.