Qwen 2 57B A14B Instruct — benchmark results

Alibaba's Apache-2.0 MoE instruct model from Qwen2 (June 2024) with 57B total and 14B active parameters and 64K context. Provider: Alibaba. Released 2024-06-04. Access: Open.

Unified ELO 1452 ± 17, rank #985 of 1776 rated models, from 78 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open Chinese LLM - GSM8K69.29Accuracy (%)94.2
Open Chinese LLM Leaderboard68.71Average Score (%)92.9
Open Chinese LLM - WinoGrande70.01Accuracy (%)91.5
Open Chinese LLM - C-Eval Semantic88.44Accuracy (%)88.1
Open Chinese LLM - HellaSwag68.48Accuracy (%)88.1
Open Chinese LLM - ARC Challenge59.47Accuracy (%)84.7
Open LLM Leaderboard - MMLU-Pro39.73Score84.7
Open Chinese LLM - CMMLU69.03Accuracy (%)81.9
Open LLM Leaderboard - GPQA10.85Score81
Open LLM Leaderboard - BBH41.79Score80.9
Open LLM Leaderboard - MATH Level 528.17Score80.4
LogicKor - Understanding9.07Score (0-10)79.9

Interactive version: theaggregate.ai/model?slug=qwen-2-57b-a14b-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.