Qwen 3 14B (Reasoning) — benchmark results

Qwen 3 14B evaluated with reasoning enabled. Provider: Alibaba. Released 2025-04-28. Access: Open.

Unified ELO 1582 ± 13, rank #526 of 1776 rated models, from 73 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Medmarks - Med-HALT Reasoning NOTA73.78Score (%)95.7
AA MATH-50096.13Accuracy (%)83.2
BRIDGE Medical Leaderboard - Zero-Shot40.17Average Performance (%)83
BRIDGE Medical Leaderboard41.34Average Performance (%)74.5
BRIDGE Medical Leaderboard - CoT37.07Average Performance (%)74.5
Medmarks - LongHealth Task 188.25Score (%)74.3
Medmarks - M-ARC45Score (%)72.9
Medmarks - SCTpublic66.36Score (%)72.9
Medmarks - HEAD-QA v285.16Score (%)67.1
Medmarks - MedConceptsQA Hard55.2Score (%)67.1
BRIDGE Medical Leaderboard - Few-Shot46.79Average Performance (%)65.1
Medmarks - Med-HALT Reasoning FCT71.25Score (%)64.3

Interactive version: theaggregate.ai/model?slug=qwen-3-14b-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.