Nanbeige4.1-3B: benchmark results

Nanbeige Lab's 3B reasoning and agentic model, an SFT+RL post-train of its Nanbeige4-3B base aimed at matching much larger Qwen3 models (February 2026). Provider: Nanbeige. Released 2026-02-11. Access: Open.

Unified ELO 1464 ± 1, rank #879 of 1392 rated models, from 37 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA GPQA Diamond84.95Accuracy (%)78
OpenEvals - GPQA Diamond83.8Accuracy (%)65.4
AA Humanity's Last Exam10.89Accuracy (%)56.7
AA Omniscience-42.5Score45.2
Artificial Analysis Intelligence Index5.34Intelligence Index35.7
AA Omniscience - Software Engineering (SWE) - Julia4Accuracy (%)30.4
BenchmarkList ECI96.02Capability Index (ECI)28.9
AA IFBench35.44Accuracy (%)27.4
AA CritPt0Accuracy (%)24.4
UGI - Willingness (W/10)3W/10 Score22.7
AA TAU-2 Bench21.64Accuracy (%)22.4
AA Omniscience - Science, Engineering & Mathematics19.2Accuracy (%)21.9

Interactive version: theaggregate.ai/model?slug=nanbeige4-1-3b · How It Works · Data refreshed daily, snapshot 2026-09-05.