Nanbeige4-3B-Thinking-2511: benchmark results

Chinese Nanbeige Lab's 3B thinking model (November 2025 update), distillation- and RL-trained for outsized function-calling and writing performance. Provider: Nanbeige. Released 2025-11-20. Access: Open.

Unified ELO 1505 ± 1, rank #670 of 1392 rated models, from 11 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
WritingBench79.03Score (self-reported)79.6
Gorilla API Bench (BFCL)51.4Overall Accuracy (%)73.2
UGI - Willingness (W/10)4.5W/10 Score34.4
UGI Leaderboard24.14UGI Score18
EQ-Bench Creative Writing v3814.6Elo17.3
Creative Writing v3965.5Elo score (self-reported)16.4
UGI - Natural Intelligence13.72NatInt Score13.5
EQ-Bench Longform Writing31.5Writing Score (0-100)6.6
EQ-Bench 31014.3Elo5.5
EQ-Bench675Normalized Elo (self-reported)5.1
UGI - Writing9.76Writing Score1.8

Interactive version: theaggregate.ai/model?slug=nanbeige4-3b-thinking-2511 · How It Works · Data refreshed daily, snapshot 2026-09-05.