Qwen 3 30B A3B (Thinking): benchmark results

Alibaba Qwen 3 30B A3B evaluated with thinking enabled. Provider: Alibaba. Released 2025-04-29. Access: Open.

Unified ELO 1565 ± 1, rank #870 of 3078 rated models, from 159 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Medmarks - PubMedQA77.47Score (%)90
Medmarks - M-ARC61.33Score (%)88.6
BRIDGE Medical Leaderboard - CoT39.35Average Performance (%)88
Medmarks - SuperGPQA Medicine Hard46.39Score (%)85.7
BRIDGE Medical Leaderboard - Zero-Shot40.93Average Performance (%)84.3
BRIDGE Medical Leaderboard42.43Average Performance (%)82.4
Medmarks - SuperGPQA Medicine Easy58.23Score (%)80.7
Medmarks - Med-HALT Reasoning NOTA67.19Score (%)80
Swallow - English MT-Bench - Reasoning86.5Judge Score (normalized, %)78.4
Medmarks - MedConceptsQA Hard59.6Score (%)75.7
Medmarks - MedHallu Medium62.58Score (%)75.7
Swallow - English MT-Bench - Math98.6Judge Score (normalized, %)75.4

Interactive version: theaggregate.ai/model?slug=qwen-3-30b-a3b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.