Qwen 3 30B A3B (Thinking) — benchmark results

Alibaba Qwen 3 30B A3B evaluated with thinking enabled. Provider: Alibaba. Released 2025-04-28. Access: Open.

Unified ELO 1590 ± 12, rank #503 of 1776 rated models, from 91 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Medmarks - PubMedQA77.47Score (%)90
BRIDGE Medical Leaderboard - CoT39.35Average Performance (%)89.6
Medmarks - M-ARC61.33Score (%)88.6
BRIDGE Medical Leaderboard - Zero-Shot40.93Average Performance (%)85.8
Medmarks - SuperGPQA Medicine Hard46.39Score (%)85.7
BRIDGE Medical Leaderboard42.43Average Performance (%)84
AA MATH-50095.93Accuracy (%)82.2
Medmarks - SuperGPQA Medicine Easy58.23Score (%)80.7
Medmarks - Med-HALT Reasoning NOTA67.19Score (%)80
Medmarks - MedConceptsQA Hard59.6Score (%)75.7
Medmarks - MedHallu Medium62.58Score (%)75.7
Medmarks - CareQA EN89.25Score (%)75

Interactive version: theaggregate.ai/model?slug=qwen-3-30b-a3b-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.