Qwen 3 Next 80B A3B (Thinking) — benchmark results

Alibaba Qwen 3 Next 80B A3B evaluated with thinking enabled. Provider: Alibaba. Released 2025-09-11. Access: Open.

Unified ELO 1614 ± 8, rank #447 of 1776 rated models, from 315 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Medmarks - Med-HALT Reasoning NOTA78.86Score (%)100
EuroEval Lithuanian Knowledge87.51Knowledge Average Score (%)99.5
Medmarks - M-ARC75Score (%)98.6
RewardBench 2 Safety94.89Accuracy (%)98
EuroEval Polish Common Sense Reasoning69.07Common Sense Reasoning Average Score (%)97.7
EuroEval Italian NLU - MultiNERD IT85.32Named entity recognition Score (%)97.4
BRIDGE Medical Leaderboard - CoT42.9Average Performance (%)97.2
EuroEval French NLU - Eltec75.31Named entity recognition Score (%)97.1
Medmarks - SuperGPQA Medicine Hard49.46Score (%)97.1
EuroEval Swedish Knowledge84.63Knowledge Average Score (%)97
EuroEval Spanish Knowledge84.37Knowledge Average Score (%)96.6
EuroEval Dutch NLU - CoNLL NL75.09Named entity recognition Score (%)96.5

Interactive version: theaggregate.ai/model?slug=qwen-3-next-80b-a3b-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.