Qwen 3 14B (Reasoning): benchmark results
Qwen 3 14B evaluated with reasoning enabled. Provider: Alibaba. Released 2025-04-29. Access: Open.
Unified ELO 1552 ± 1, rank #954 of 3078 rated models, from 252 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Nejumi 4 - jaster (2-shot) - JSICK | 86 | Exact match (%) | 97.1 |
| Medmarks - Med-HALT Reasoning NOTA | 73.78 | Score (%) | 95.7 |
| Avalon-ToM-Bench - Strategic Signaling | 80.71 | Accuracy (%) | 88.5 |
| Nejumi 4 - BFCL - Non-Live AST | 80.74 | Accuracy (%) | 87.1 |
| Nejumi 4 - jaster (0-shot) - JSICK | 83 | Exact match (%) | 84.6 |
| Nejumi 4 - BFCL - Relevance Detection | 83.33 | Accuracy (%) | 81.6 |
| BRIDGE Medical Leaderboard - Zero-Shot | 40.17 | Average Performance (%) | 81.5 |
| Nejumi 4 - BFCL - Live AST | 72.22 | Accuracy (%) | 76.8 |
| Nejumi 4 - Toxicity - Fairness | 97.74 | Criteria met (%) | 76.1 |
| Swallow - Japanese MT-Bench - Reasoning | 78.3 | Judge Score (normalized, %) | 76.1 |
| Swallow - English MT-Bench - Math | 98.6 | Judge Score (normalized, %) | 75.4 |
| Swallow - English MT-Bench - Extraction | 80.4 | Judge Score (normalized, %) | 74.6 |
Interactive version: theaggregate.ai/model?slug=qwen-3-14b-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-19.