Nanbeige4.1-3B — benchmark results

Nanbeige Lab's 3B reasoning and agentic model, an SFT+RL post-train of its Nanbeige4-3B base aimed at matching much larger Qwen3 models (February 2026). Provider: Nanbeige. Released 2026-02-11. Access: Open.

Unified ELO 1452 ± 34, rank #986 of 1776 rated models, from 42 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA GPQA Diamond84.95Accuracy (%)83.9
OpenEvals - GPQA Diamond83.8Accuracy (%)65.4
AA Humanity's Last Exam10.01Accuracy (%)61.8
AA Omniscience-41.93Score50.9
Artificial Analysis Intelligence Index11.09Intelligence Index38.4
AA SciCode26.6Accuracy (%)33
CritPt0Accuracy (self-reported)30.3
AA IFBench35.44Accuracy (%)27.5
AA CritPt0Accuracy (%)27.1
AA Omniscience - Science, Engineering & Mathematics19Accuracy (%)26.3
AA Omniscience - Software Engineering (SWE) - Julia4Accuracy (%)23.7
AA TAU-2 Bench21.64Accuracy (%)22.3

Interactive version: theaggregate.ai/model?slug=nanbeige4-1-3b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.