Qwen 3 14B (Reasoning) — benchmark results
Qwen 3 14B evaluated with reasoning enabled. Provider: Alibaba. Released 2025-04-28. Access: Open.
Unified ELO 1582 ± 13, rank #526 of 1776 rated models, from 73 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Medmarks - Med-HALT Reasoning NOTA | 73.78 | Score (%) | 95.7 |
| AA MATH-500 | 96.13 | Accuracy (%) | 83.2 |
| BRIDGE Medical Leaderboard - Zero-Shot | 40.17 | Average Performance (%) | 83 |
| BRIDGE Medical Leaderboard | 41.34 | Average Performance (%) | 74.5 |
| BRIDGE Medical Leaderboard - CoT | 37.07 | Average Performance (%) | 74.5 |
| Medmarks - LongHealth Task 1 | 88.25 | Score (%) | 74.3 |
| Medmarks - M-ARC | 45 | Score (%) | 72.9 |
| Medmarks - SCTpublic | 66.36 | Score (%) | 72.9 |
| Medmarks - HEAD-QA v2 | 85.16 | Score (%) | 67.1 |
| Medmarks - MedConceptsQA Hard | 55.2 | Score (%) | 67.1 |
| BRIDGE Medical Leaderboard - Few-Shot | 46.79 | Average Performance (%) | 65.1 |
| Medmarks - Med-HALT Reasoning FCT | 71.25 | Score (%) | 64.3 |
Interactive version: theaggregate.ai/model?slug=qwen-3-14b-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.