Qwen 3 235B A22B (Thinking): benchmark results
Alibaba Qwen 3 235B A22B evaluated with thinking enabled. Provider: Alibaba. Released 2025-04-28. Access: Open.
Unified ELO 1592 ± 1, rank #469 of 1761 rated models, from 96 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Medmarks - Med-HALT Reasoning FCT | 90.07 | Score (%) | 95.7 |
| Medmarks - LongHealth Task 1 | 90.67 | Score (%) | 94.3 |
| Medmarks - LongHealth Task 2 | 90.25 | Score (%) | 94.3 |
| Medmarks - MedQA | 92.96 | Score (%) | 94.3 |
| Medmarks - SuperGPQA Medicine Easy | 66.7 | Score (%) | 94.3 |
| Medmarks - M-ARC | 66 | Score (%) | 92.9 |
| Medmarks - Medbullets OP4 | 85.71 | Score (%) | 92.9 |
| Medmarks - CareQA EN | 93.13 | Score (%) | 91.4 |
| Medmarks - HEAD-QA v2 | 90.05 | Score (%) | 91.4 |
| Medmarks - MedConceptsQA Easy | 99.47 | Score (%) | 91.4 |
| Medmarks - MedConceptsQA Hard | 79.27 | Score (%) | 91.4 |
| Medmarks - MedConceptsQA Medium | 85.95 | Score (%) | 91.4 |
Interactive version: theaggregate.ai/model?slug=qwen-3-235b-a22b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.