Qwen-7B — benchmark results
Provider: Alibaba. Released 2023-08-01. Access: Open.
Unified ELO 1302 ± 14, rank #1531 of 1776 rated models, from 28 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ARC Challenge (AI2) | 75.3 | Accuracy (%) | 70.5 |
| T-Eval | 59.5 | Overall Score (%) | 70 |
| GSM8K | 51.7 | Accuracy (%) | 52.1 |
| CyberMetric | 52.9 | Accuracy (%) | 41.7 |
| CMMLU | 58.66 | 5-shot Avg Accuracy (%) | 40 |
| Big-Bench Hard | 45 | Average (%) | 39.6 |
| BoolQ | 76.4 | Accuracy (%) | 38.3 |
| VMLU | 32.81 | Average (%) | 37.5 |
| VMLU - Humanities | 34.15 | Accuracy (%) | 37.5 |
| VMLU - Other | 32.68 | Accuracy (%) | 37.5 |
| VMLU - STEM | 30.64 | Accuracy (%) | 37.5 |
| InfiBench | 31.69 | Score (%) | 35.2 |
Interactive version: theaggregate.ai/model?slug=qwen-7b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.