Qwen 3 Next 80B A3B (Thinking): benchmark results

Alibaba Qwen 3 Next 80B A3B evaluated with thinking enabled. Provider: Alibaba. Released 2025-09-11. Access: Open.

Unified ELO 1562 ± 1, rank #591 of 1761 rated models, from 403 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
MERA Code - ruCodeEval76.22pass@1 (%)100
MERA Code - ruHumanEval82.26pass@1 (%)100
Medmarks - Med-HALT Reasoning NOTA78.86Score (%)100
EuroEval Lithuanian Knowledge87.51Knowledge Average Score (%)99.5
MERA - ruMultiAr100EM (%)98.6
Medmarks - M-ARC75Score (%)98.6
RewardBench 2 Safety94.89Accuracy (%)98
EuroEval Polish Common Sense Reasoning69.07Common Sense Reasoning Average Score (%)97.7
EuroEval Italian NLU - MultiNERD IT85.32Named entity recognition Score (%)97.4
EuroEval French NLU - Eltec75.31Named entity recognition Score (%)97.1
Medmarks - SuperGPQA Medicine Hard49.46Score (%)97.1
EuroEval Swedish Knowledge84.63Knowledge Average Score (%)97

Interactive version: theaggregate.ai/model?slug=qwen-3-next-80b-a3b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.