Qwen 3 4B (Reasoning): benchmark results
Qwen 3 4B evaluated with reasoning enabled. Provider: Alibaba. Released 2025-04-29. Access: Open.
Unified ELO 1494 ± 1, rank #1511 of 3078 rated models, from 106 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Medmarks - M-ARC | 57.67 | Score (%) | 81.4 |
| Avalon-ToM-Bench - Intention Attribution | 86.27 | Accuracy (%) | 76.9 |
| Medmarks - MedHallu Easy | 67.89 | Score (%) | 75.7 |
| Medmarks - MedHallu Hard | 47.93 | Score (%) | 75.7 |
| Medmarks - MedHallu Medium | 62.12 | Score (%) | 72.9 |
| BRIDGE Medical Leaderboard - CoT | 36.98 | Average Performance (%) | 70.4 |
| Swallow - Japanese MT-Bench - Extraction | 73.4 | Judge Score (normalized, %) | 69.4 |
| Avalon-ToM-Bench - Strategic Signaling | 76.87 | Accuracy (%) | 69.2 |
| Swallow - English MT-Bench - Reasoning | 84.1 | Judge Score (normalized, %) | 67.2 |
| BRIDGE Medical Leaderboard - Zero-Shot | 38.5 | Average Performance (%) | 65.7 |
| Swallow - English MT-Bench - Math | 98.3 | Judge Score (normalized, %) | 65.7 |
| Medmarks - Med-HALT Reasoning NOTA | 59.38 | Score (%) | 62.9 |
Interactive version: theaggregate.ai/model?slug=qwen-3-4b-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-19.