Qwen 3 4B — benchmark results

Alibaba's 4B dense Qwen3 model with switchable thinking/non-thinking modes, matching prior 7B-class Qwen2.5 quality (April 2025). Provider: Alibaba. Released 2025-04-28. Access: Open.

Unified ELO 1436 ± 8, rank #1060 of 1776 rated models, from 320 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
ChineseSafe Benchmark74.95Accuracy (%)92.6
BTZSC64.86Macro-F1 (%)91.2
BTZSC - Topic63.83Macro-F1 (%)91.2
FACTS Leaderboard42.55Combined Score (%)90.9
SLMJury89.2Judge accuracy (%)86.7
BTZSC - Sentiment88.32Macro-F1 (%)85.3
LA Leaderboard - Spanish Law Exams38.66Accuracy (%)83.1
Vectara Hallucination Leaderboard94.3Factual Consistency Rate (%)79.8
EuroEval Spanish NLU - Sentiment Headlines ES47.57Sentiment classification Score (%)78.9
Open-R1 Eval Leaderboard65.61Average Accuracy (%)77.8
EuroEval Spanish NLU50.8NLU Average Score (%)77.6
EuroEval Portuguese Knowledge67.43Knowledge Average Score (%)77.3

Interactive version: theaggregate.ai/model?slug=qwen-3-4b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.