Qwen 3 8B (Thinking): benchmark results

Qwen 3 8B evaluated with thinking enabled. Provider: Alibaba. Released 2025-04-29. Access: Open.

Unified ELO 1521 ± 1, rank #1203 of 3078 rated models, from 233 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
GeoBenchLLM - GKMC82Accuracy (%)100
GeoBenchLLM - GeoSQA63Accuracy (%)100
GeoBenchLLM - GridRoute81Optimal route ratio (%)100
GeoBenchLLM - PPNL Multi-Goal62Optimal route ratio (%)100
GeoBenchLLM - PPNL Single-Goal92Optimal route ratio (%)100
LLM Self-Modeling - Confidence-Recall-0.08Skill above dummy predictor (-1 to 1, higher is better; 4096100
LLM Self-Modeling - Edit-Proposal0.57Skill above dummy predictor (-1 to 1, higher is better; 4096100
LLM Self-Modeling - Feature-Rate0.05Skill above dummy predictor (-1 to 1, higher is better; 4096100
LLM Self-Modeling - Output-Prediction0.49Skill above dummy predictor (-1 to 1, higher is better; 4096100
LLM Self-Modeling - Perturbation-Choice0.07Skill above dummy predictor (-1 to 1, higher is better; 4096100
LLM Self-Modeling0.11Skill above dummy predictor (-1 to 1, higher is better; 409683.3
Medmarks - Med-HALT Reasoning NOTA67.8Score (%)82.9

Interactive version: theaggregate.ai/model?slug=qwen-3-8b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.