Qwen 3 235B A22B (Thinking) — benchmark results

Alibaba Qwen 3 235B A22B evaluated with thinking enabled. Provider: Alibaba. Released 2025-04-28. Access: Open.

Unified ELO 1629 ± 9, rank #416 of 1776 rated models, from 206 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EsoBench29.4Score100
LLM Stats (Multi-IF)80.6Score (%)100
LLM Stats (WritingBench)88.3Score (%)100
Medmarks - Med-HALT Reasoning FCT90.07Score (%)95.7
Medmarks - LongHealth Task 190.67Score (%)94.3
Medmarks - LongHealth Task 290.25Score (%)94.3
Medmarks - MedQA92.96Score (%)94.3
Medmarks - SuperGPQA Medicine Easy66.7Score (%)94.3
LongBench v260.6Accuracy (%)94.1
Medmarks - M-ARC66Score (%)92.9
Medmarks - Medbullets OP485.71Score (%)92.9
Medmarks - CareQA EN93.13Score (%)91.4

Interactive version: theaggregate.ai/model?slug=qwen-3-235b-a22b-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.