MultiNRC — leaderboard

MultiNRC benchmarks LLMs on 1,000+ culturally grounded reasoning questions by native French, Spanish, and Chinese speakers across four reasoning categor...

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 39 models tracked.

Top models

#ModelScore
1GPT-5 Pro65.2
2Gemini 3.1 Pro (Preview)64.74
3GPT-5.4 Pro (xHigh)62.27
4Muse Spark59.05
5GPT-5.4 (xHigh)58.29
6Claude Opus 4.6 (Max)57.06
7GPT-552.13
8GPT-5.149
9Claude Opus 4.548.63
10O3 (2025-04-16) (High)45.5
11Gemini 2.5 Pro (Preview 06-05)45.12
12O3 (2025-04-16) (Medium)44.45
13GPT-5.242.18
14Claude Opus 4.138.39
15Claude Sonnet 4.535.83

Interactive version: theaggregate.ai/benchmark?slug=multinrc · How the rankings work · Data refreshed daily, snapshot 2026-07-22.