Qwen 1.5 110B — benchmark results
Alibaba's first 100B+ open-weights Qwen release (April 2024), a 110B dense base model with GQA and 32K context. Provider: Alibaba. Released 2024-04-25. Access: Open.
Unified ELO 1491 ± 23, rank #833 of 1776 rated models, from 18 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open LLM Leaderboard - MMLU-Pro | 48.45 | Score | 94 |
| Open LLM Leaderboard - GPQA | 13.65 | Score | 88.6 |
| CMMLU | 88.32 | 5-shot Avg Accuracy (%) | 87.5 |
| Open LLM Leaderboard - BBH | 44.28 | Score | 84.6 |
| Open LLM Leaderboard - MATH Level 5 | 24.7 | Score | 77.8 |
| Open LLM Leaderboard - MuSR | 13.71 | Score | 75.4 |
| HELM NaturalQuestions (Open) | 73.95 | F1 (%) | 70 |
| HELM Lite | 59.84 | Mean win rate (self-reported) | 67.5 |
| SynthPAI | 65.7 | Average accuracy in % | 64.7 |
| HELM (Stanford) | 54.96 | Mean Win Rate (%) | 60 |
| ASCIIEval | 30.28 | Macro-Avg Accuracy (%) | 52.8 |
| HELM WMT 2014 | 19.24 | BLEU-4 (%) | 52.2 |
Interactive version: theaggregate.ai/model?slug=qwen-1-5-110b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.