Qwen 3 235B A22B 2507 (Thinking): benchmark results

Qwen 3 235B A22B 2507 evaluated with thinking enabled. Provider: Alibaba. Released 2025-07-01. Access: Open.

Unified ELO 1581 ± 1, rank #508 of 1761 rated models, from 241 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EsoBench29.4Score100
LLM Stats (Multi-IF)80.6Score (%)100
LLM Stats (WritingBench)88.3Score (%)100
MERA - MathLogicQA99.74Accuracy (%)98.6
MERA - RCB61.19Accuracy (%)98.3
MERA - ruHateSpeech92.83Accuracy (%)97.6
MERA - ruTiE95.02Accuracy (%)97.1
MERA0.8Total score96.4
MERA - ruMultiAr99.9EM (%)96.4
WritingBench82.34Score (self-reported)96.3
MERA - ruCodeEval71.89pass@1 (%)96.2
MERA - USE68.63Grade, normalized (%)95.9

Interactive version: theaggregate.ai/model?slug=qwen-3-235b-a22b-2507-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.