Qwen 3 30B A3B 2507 (Thinking) — benchmark results
Qwen 3 30B A3B 2507 thinking snapshot. Provider: Alibaba. Released 2025-07-01. Access: Open.
Unified ELO 1583 ± 18, rank #521 of 1776 rated models, from 106 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HealthBench Hard | 58 | Overall score (self-reported) | 96.3 |
| BRIDGE Medical Leaderboard - CoT | 41.43 | Average Performance (%) | 94.3 |
| BRIDGE Medical Leaderboard | 44.5 | Average Performance (%) | 92.5 |
| RewardBench 2 Safety | 93.61 | Accuracy (%) | 92.2 |
| BRIDGE Medical Leaderboard - Zero-Shot | 41.82 | Average Performance (%) | 91.5 |
| AA MATH-500 | 97.6 | Accuracy (%) | 89.1 |
| BRIDGE Medical Leaderboard - Few-Shot | 50.26 | Average Performance (%) | 85.8 |
| MATH-MC Level 5 | 99.61 | Accuracy (%) | 85.3 |
| K-MetBench | 76.7 | Accuracy (self-reported) | 84.5 |
| AA LiveCodeBench | 70.69 | Pass@1 (%) | 80.8 |
| LLM2014 Logic 2025-09 | 46.52 | Median Score | 75.6 |
| AA-LCR | 59 | Score (self-reported) | 73.9 |
Interactive version: theaggregate.ai/model?slug=qwen-3-30b-a3b-2507-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.