Qwen 3 14B (Non-reasoning): benchmark results
Qwen 3 14B evaluated with reasoning disabled. Provider: Alibaba. Released 2025-04-28. Access: Open.
Unified ELO 1495 ± 1, rank #952 of 1919 rated models, from 288 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| When Simulation Lies | 52.92 | Pert Acc (self-reported) | 100 |
| EuroEval Dutch Summarization - Wiki Lingua NL | 36.05 | Score (%) | 98.7 |
| EuroEval Latvian NLU - MultiWikiQA LV | 72.18 | Reading comprehension Score (%) | 98.7 |
| EuroEval Danish NLU - Multi Wiki QA DA | 82.01 | Reading comprehension Score (%) | 97.8 |
| EuroEval Albanian NLU - MultiWikiQA SQ | 68.4 | Reading comprehension Score (%) | 97.5 |
| EuroEval Portuguese NLU - MultiWikiQA PT | 78.55 | Reading comprehension Score (%) | 97.2 |
| EuroEval Slovene NLU - MultiWikiQA SL | 72.27 | Reading comprehension Score (%) | 96.6 |
| EuroEval Icelandic Summarization - RRN | 38.85 | Score (%) | 96.1 |
| EuroEval Catalan NLU - MultiWikiQA CA | 76.65 | Reading comprehension Score (%) | 95.7 |
| EuroEval French NLU - FQuAD | 75.59 | Reading comprehension Score (%) | 95.7 |
| EuroEval Dutch Simplification - Duidelijke Taal | 56.34 | Score (%) | 95.5 |
| EuroEval Hungarian NLU - HuSST | 62.4 | Sentiment classification Score (%) | 95.5 |
Interactive version: theaggregate.ai/model?slug=qwen-3-14b-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-08.