Qwen 3 4B (Reasoning): benchmark results

Qwen 3 4B evaluated with reasoning enabled. Provider: Alibaba. Released 2025-04-29. Access: Open.

Unified ELO 1494 ± 1, rank #1511 of 3078 rated models, from 106 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Medmarks - M-ARC57.67Score (%)81.4
Avalon-ToM-Bench - Intention Attribution86.27Accuracy (%)76.9
Medmarks - MedHallu Easy67.89Score (%)75.7
Medmarks - MedHallu Hard47.93Score (%)75.7
Medmarks - MedHallu Medium62.12Score (%)72.9
BRIDGE Medical Leaderboard - CoT36.98Average Performance (%)70.4
Swallow - Japanese MT-Bench - Extraction73.4Judge Score (normalized, %)69.4
Avalon-ToM-Bench - Strategic Signaling76.87Accuracy (%)69.2
Swallow - English MT-Bench - Reasoning84.1Judge Score (normalized, %)67.2
BRIDGE Medical Leaderboard - Zero-Shot38.5Average Performance (%)65.7
Swallow - English MT-Bench - Math98.3Judge Score (normalized, %)65.7
Medmarks - Med-HALT Reasoning NOTA59.38Score (%)62.9

Interactive version: theaggregate.ai/model?slug=qwen-3-4b-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-19.