NonoBench (v1.0): leaderboard

Metric: Standard accuracy (%), correct / non-failed attempts; version 1.0. Source: www.nonobench.com. Saturation forecast: Around January 2027. 38 models tracked.

Top models

#ModelScore
1GPT-5.2 (High)66.67
2Claude Opus 4.5 (High)60
3Gemini 3 Pro (Preview) (High)56.67
4GPT-5.2 (xHigh)56.67
5Gemini 3 Flash (Preview) (High)50
6GPT-5.2 (Low)50
7GPT-OSS-120B (High)43.33
8DeepSeek V3.2 (High)43.33
9Kimi K2.5 (High)43.33
10DeepSeek V3.2 Speciale40
11Grok 433.33
12Qwen 3 Next 80B A3B (Thinking)33.33
13GPT-OSS-120B (Low)33.33
14Claude Sonnet 4.5 (Thinking)30
15Kimi K2 (Thinking)26.67

Interactive version: theaggregate.ai/benchmark?slug=nonobench-v1-0 · How It Works · Data refreshed daily, snapshot 2026-10-09.