Qwen 2.5 3B — benchmark results
Alibaba's 3B model from the Qwen2.5 series (September 2024), pretrained on up to 18T tokens and released under the non-commercial Qwen Research License. Provider: Alibaba. Released 2024-09-19. Access: Open.
Unified ELO 1408 ± 14, rank #1183 of 1776 rated models, from 211 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Relevance as a Vulnerability | 3.22 | Agent (self-reported) | 85.7 |
| Open Arabic LLM - Arabic MMLU Civics (High School) | 50.57 | Accuracy (%) | 80.6 |
| Open Japanese LLM - Mbpp Pylint Check | 90.76 | Score (%) | 77.9 |
| LA Leaderboard - AQuAS | 67.79 | Accuracy (%) | 76.5 |
| LA Leaderboard - GalCoLA | 53.42 | Accuracy (%) | 73.5 |
| LA Leaderboard - XNLI Galician | 46.8 | Accuracy (%) | 73.5 |
| LA Leaderboard | 55.69 | Average Score (%) | 72.1 |
| LA Leaderboard - Spanish Law Exams | 34.45 | Accuracy (%) | 70.6 |
| Open Japanese LLM - CG | 45.78 | Score (%) | 68.1 |
| Open Arabic LLM - Alghafa Multiple Choice Rating Sentiment Task | 52.64 | Accuracy (%) | 65.4 |
| Open Arabic LLM - Alghafa Multiple Choice Sentiment Task | 39.65 | Accuracy (%) | 63.9 |
| Open Japanese LLM - Xlsum JA Rouge1 | 27.34 | Score (%) | 63.2 |
Interactive version: theaggregate.ai/model?slug=qwen-2-5-3b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.