Qwen 3 14B — benchmark results

Alibaba Qwen 3 14B model row. Provider: Alibaba. Released 2025-04-28. Access: Open.

Unified ELO 1454 ± 14, rank #977 of 1776 rated models, from 444 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AgingBench7.9S1 kw_m HL (self-reported)100
AppWorld Challenge67.6Task Goal Completion (%)100
AppWorld Normal86.9Task Goal Completion (%)100
ChLogic99.13English (self-reported)100
TSCG90.220 Tools (json-text) (self-reported)100
When Simulation Lies52.9Pert Acc (self-reported)100
AGC-Bench - unfun_corpus1.44Dataset z-score98.7
INCLUDE-base-44 European Languages63.03Average Accuracy (%)97.1
EuroEval German NLU - GermEval73.49Named entity recognition Score (%)95.7
EuroEval Lithuanian58.25Average Score (%)95.2
EuroEval Lithuanian NLU55.05NLU Average Score (%)95.2
AI Chess Leaderboard (Continuation)1409Elo94.7

Interactive version: theaggregate.ai/model?slug=qwen-3-14b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.