Qwen 3 8B (Thinking) — benchmark results

Qwen 3 8B evaluated with thinking enabled. Provider: Alibaba. Released 2025-04-28. Access: Open.

Unified ELO 1532 ± 12, rank #680 of 1776 rated models, from 93 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Medmarks - Med-HALT Reasoning NOTA67.8Score (%)82.9
BRIDGE Medical Leaderboard - Zero-Shot39.98Average Performance (%)81.1
Medmarks - SCTpublic67.82Score (%)80
BRIDGE Medical Leaderboard - CoT37.06Average Performance (%)73.6
BRIDGE Medical Leaderboard40.74Average Performance (%)70.8
Medmarks - LongHealth Task 287.88Score (%)66.4
AA MATH-50090.4Accuracy (%)66.3
Medmarks - LongHealth Task 187.17Score (%)58.6
BRIDGE Medical Leaderboard - Few-Shot45.19Average Performance (%)58.5
AA Omniscience - Software Engineering (SWE) - Dart22Accuracy (%)54.7
Medmarks - HEAD-QA v282.43Score (%)52.1
Medmarks - PubMedQA74.8Score (%)52.1

Interactive version: theaggregate.ai/model?slug=qwen-3-8b-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.