Qwen 3 8B (Thinking): benchmark results
Qwen 3 8B evaluated with thinking enabled. Provider: Alibaba. Released 2025-04-29. Access: Open.
Unified ELO 1521 ± 1, rank #1203 of 3078 rated models, from 233 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| GeoBenchLLM - GKMC | 82 | Accuracy (%) | 100 |
| GeoBenchLLM - GeoSQA | 63 | Accuracy (%) | 100 |
| GeoBenchLLM - GridRoute | 81 | Optimal route ratio (%) | 100 |
| GeoBenchLLM - PPNL Multi-Goal | 62 | Optimal route ratio (%) | 100 |
| GeoBenchLLM - PPNL Single-Goal | 92 | Optimal route ratio (%) | 100 |
| LLM Self-Modeling - Confidence-Recall | -0.08 | Skill above dummy predictor (-1 to 1, higher is better; 4096 | 100 |
| LLM Self-Modeling - Edit-Proposal | 0.57 | Skill above dummy predictor (-1 to 1, higher is better; 4096 | 100 |
| LLM Self-Modeling - Feature-Rate | 0.05 | Skill above dummy predictor (-1 to 1, higher is better; 4096 | 100 |
| LLM Self-Modeling - Output-Prediction | 0.49 | Skill above dummy predictor (-1 to 1, higher is better; 4096 | 100 |
| LLM Self-Modeling - Perturbation-Choice | 0.07 | Skill above dummy predictor (-1 to 1, higher is better; 4096 | 100 |
| LLM Self-Modeling | 0.11 | Skill above dummy predictor (-1 to 1, higher is better; 4096 | 83.3 |
| Medmarks - Med-HALT Reasoning NOTA | 67.8 | Score (%) | 82.9 |
Interactive version: theaggregate.ai/model?slug=qwen-3-8b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.