Qwen 3 4B — benchmark results
Alibaba's 4B dense Qwen3 model with switchable thinking/non-thinking modes, matching prior 7B-class Qwen2.5 quality (April 2025). Provider: Alibaba. Released 2025-04-28. Access: Open.
Unified ELO 1436 ± 8, rank #1060 of 1776 rated models, from 320 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ChineseSafe Benchmark | 74.95 | Accuracy (%) | 92.6 |
| BTZSC | 64.86 | Macro-F1 (%) | 91.2 |
| BTZSC - Topic | 63.83 | Macro-F1 (%) | 91.2 |
| FACTS Leaderboard | 42.55 | Combined Score (%) | 90.9 |
| SLMJury | 89.2 | Judge accuracy (%) | 86.7 |
| BTZSC - Sentiment | 88.32 | Macro-F1 (%) | 85.3 |
| LA Leaderboard - Spanish Law Exams | 38.66 | Accuracy (%) | 83.1 |
| Vectara Hallucination Leaderboard | 94.3 | Factual Consistency Rate (%) | 79.8 |
| EuroEval Spanish NLU - Sentiment Headlines ES | 47.57 | Sentiment classification Score (%) | 78.9 |
| Open-R1 Eval Leaderboard | 65.61 | Average Accuracy (%) | 77.8 |
| EuroEval Spanish NLU | 50.8 | NLU Average Score (%) | 77.6 |
| EuroEval Portuguese Knowledge | 67.43 | Knowledge Average Score (%) | 77.3 |
Interactive version: theaggregate.ai/model?slug=qwen-3-4b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.