Qwen 3 235B A22B (Thinking) — benchmark results
Alibaba Qwen 3 235B A22B evaluated with thinking enabled. Provider: Alibaba. Released 2025-04-28. Access: Open.
Unified ELO 1629 ± 9, rank #416 of 1776 rated models, from 206 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EsoBench | 29.4 | Score | 100 |
| LLM Stats (Multi-IF) | 80.6 | Score (%) | 100 |
| LLM Stats (WritingBench) | 88.3 | Score (%) | 100 |
| Medmarks - Med-HALT Reasoning FCT | 90.07 | Score (%) | 95.7 |
| Medmarks - LongHealth Task 1 | 90.67 | Score (%) | 94.3 |
| Medmarks - LongHealth Task 2 | 90.25 | Score (%) | 94.3 |
| Medmarks - MedQA | 92.96 | Score (%) | 94.3 |
| Medmarks - SuperGPQA Medicine Easy | 66.7 | Score (%) | 94.3 |
| LongBench v2 | 60.6 | Accuracy (%) | 94.1 |
| Medmarks - M-ARC | 66 | Score (%) | 92.9 |
| Medmarks - Medbullets OP4 | 85.71 | Score (%) | 92.9 |
| Medmarks - CareQA EN | 93.13 | Score (%) | 91.4 |
Interactive version: theaggregate.ai/model?slug=qwen-3-235b-a22b-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.