Qwen 3 30B A3B 2507 (Thinking) — benchmark results

Qwen 3 30B A3B 2507 thinking snapshot. Provider: Alibaba. Released 2025-07-01. Access: Open.

Unified ELO 1583 ± 18, rank #521 of 1776 rated models, from 106 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HealthBench Hard58Overall score (self-reported)96.3
BRIDGE Medical Leaderboard - CoT41.43Average Performance (%)94.3
BRIDGE Medical Leaderboard44.5Average Performance (%)92.5
RewardBench 2 Safety93.61Accuracy (%)92.2
BRIDGE Medical Leaderboard - Zero-Shot41.82Average Performance (%)91.5
AA MATH-50097.6Accuracy (%)89.1
BRIDGE Medical Leaderboard - Few-Shot50.26Average Performance (%)85.8
MATH-MC Level 599.61Accuracy (%)85.3
K-MetBench76.7Accuracy (self-reported)84.5
AA LiveCodeBench70.69Pass@1 (%)80.8
LLM2014 Logic 2025-0946.52Median Score75.6
AA-LCR59Score (self-reported)73.9

Interactive version: theaggregate.ai/model?slug=qwen-3-30b-a3b-2507-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.