Qwen 3 8B (Non-reasoning): benchmark results
Qwen 3 8B evaluated with reasoning disabled. Provider: Alibaba. Released 2025-04-29. Access: Open.
Unified ELO 1479 ± 1, rank #1664 of 3078 rated models, from 350 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EuroEval Ukrainian Summarization - LR SUM UK | 31.21 | Score (%) | 93.8 |
| EuroEval Italian Summarization - Ilpost SUM | 38.55 | Score (%) | 93.3 |
| EuroEval Catalan Summarization - Dacsa CA | 38.65 | Score (%) | 92.2 |
| TRACE Bench - Length | 98.32 | Reply-length compliance (%) | 92 |
| EuroEval Spanish Summarization - Mlsum ES | 29.13 | Score (%) | 91.8 |
| EuroEval Ukrainian NLU - Cross Domain UK Reviews | 62.26 | Sentiment classification Score (%) | 91.8 |
| EuroEval Icelandic Summarization - RRN | 37.96 | Score (%) | 91.6 |
| EuroEval Catalan NLU - MultiWikiQA CA | 74.17 | Reading comprehension Score (%) | 89.6 |
| EuroEval Albanian NLU - MultiWikiQA SQ | 63.62 | Reading comprehension Score (%) | 89.4 |
| EuroEval Slovene NLU - MultiWikiQA SL | 68.43 | Reading comprehension Score (%) | 89.2 |
| EuroEval Polish NLU - Polemo2 | 92.97 | Sentiment classification Score (%) | 88.2 |
| EuroEval Estonian Summarization - ERR News | 30.97 | Score (%) | 87.9 |
Interactive version: theaggregate.ai/model?slug=qwen-3-8b-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-19.