Qwen 3 14B (Reasoning): benchmark results

Qwen 3 14B evaluated with reasoning enabled. Provider: Alibaba. Released 2025-04-29. Access: Open.

Unified ELO 1552 ± 1, rank #954 of 3078 rated models, from 252 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Nejumi 4 - jaster (2-shot) - JSICK86Exact match (%)97.1
Medmarks - Med-HALT Reasoning NOTA73.78Score (%)95.7
Avalon-ToM-Bench - Strategic Signaling80.71Accuracy (%)88.5
Nejumi 4 - BFCL - Non-Live AST80.74Accuracy (%)87.1
Nejumi 4 - jaster (0-shot) - JSICK83Exact match (%)84.6
Nejumi 4 - BFCL - Relevance Detection83.33Accuracy (%)81.6
BRIDGE Medical Leaderboard - Zero-Shot40.17Average Performance (%)81.5
Nejumi 4 - BFCL - Live AST72.22Accuracy (%)76.8
Nejumi 4 - Toxicity - Fairness97.74Criteria met (%)76.1
Swallow - Japanese MT-Bench - Reasoning78.3Judge Score (normalized, %)76.1
Swallow - English MT-Bench - Math98.6Judge Score (normalized, %)75.4
Swallow - English MT-Bench - Extraction80.4Judge Score (normalized, %)74.6

Interactive version: theaggregate.ai/model?slug=qwen-3-14b-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-19.