Qwen 3 4B (Reasoning) — benchmark results

Qwen 3 4B evaluated with reasoning enabled. Provider: Alibaba. Released 2025-04-28. Access: Open.

Unified ELO 1518 ± 27, rank #737 of 1776 rated models, from 44 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Medmarks - M-ARC57.67Score (%)81.4
Medmarks - MedHallu Easy67.89Score (%)75.7
Medmarks - MedHallu Hard47.93Score (%)75.7
AA MATH-50093.33Accuracy (%)74.3
Medmarks - MedHallu Medium62.12Score (%)72.9
BRIDGE Medical Leaderboard - CoT36.98Average Performance (%)71.7
BRIDGE Medical Leaderboard - Zero-Shot38.5Average Performance (%)67
Medmarks - Med-HALT Reasoning NOTA59.38Score (%)62.9
BRIDGE Medical Leaderboard39.39Average Performance (%)60.4
Medmarks - LongHealth Task 186.75Score (%)57.1
Medmarks - MedConceptsQA Hard50.8Score (%)52.9
AA LiveCodeBench46.46Pass@1 (%)52.8

Interactive version: theaggregate.ai/model?slug=qwen-3-4b-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.