Qwen 2 7B Instruct — benchmark results

Alibaba's Apache-2.0 7B Qwen2 instruct model (June 2024) handling context lengths up to 128K tokens. Provider: Alibaba. Released 2024-06-07. Access: Open.

Unified ELO 1439 ± 9, rank #1045 of 1776 rated models, from 320 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SeaEval - Multilingual Reasoning - ZBench (Zero-Shot)72.73Accuracy (%)100
SeaEval - Cultural Reasoning - SG-Eval v2 Open (Zero-Shot)56.56Accuracy (%)95.7
SeaEval - Fundamental NLP Tasks - MRPC (Zero-Shot)78.68Accuracy (%)91.3
OpenEval - BBQ91.39Exact Match (%)90.5
SeaEval - Cultural Reasoning - CN-Eval (Zero-Shot)82.86Accuracy (%)87.5
Open Arabic LLM - Arabic MMLU Computer Science (Middle School)92.59Accuracy (%)87
Open Arabic LLM - Alghafa Multiple Choice Rating Sentiment NO Neutral Task80.08Accuracy (%)82.7
LiveBench LCB Generation40Score82.6
Open CoT - LogiQA6.55CoT Gain (%)82.1
Open Arabic LLM - Aratrust Unfairness92.73Accuracy (%)82
SeaEval - Multilingual Reasoning - CMMLU (Zero-Shot)77.28Accuracy (%)81
European LLM Leaderboard - Zero-Shot Accuracy50.71Average Accuracy (%)80.5

Interactive version: theaggregate.ai/model?slug=qwen-2-7b-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.