Nanbeige4-3B-Thinking-2511 — benchmark results

Chinese Nanbeige Lab's 3B thinking model (November 2025 update), distillation- and RL-trained for outsized function-calling and writing performance. Provider: Nanbeige. Released 2025-11-01. Access: Open.

Unified ELO 1496 ± 44, rank #815 of 1776 rated models, from 11 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
WritingBench79.03Score (self-reported)80.8
Gorilla API Bench (BFCL)51.4Overall Accuracy (%)73.2
UGI - Willingness (W/10)4.5W/10 Score33.7
UGI Leaderboard24.14UGI Score18.1
Creative Writing v3953.5Elo score (self-reported)17.8
EQ-Bench Creative Writing v3814.6Elo17.3
UGI - Natural Intelligence13.72NatInt Score13.6
EQ-Bench Longform Writing31.5Writing Score (0-100)7
EQ-Bench 31014.3Elo5.5
EQ-Bench672.8Normalized Elo (self-reported)5.3
UGI - Writing9.76Writing Score1.8

Interactive version: theaggregate.ai/model?slug=nanbeige4-3b-thinking-2511 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.