OR-Space: leaderboard

Metric: Build Pass@1 (self-reported). Source: benchmarklist.com. 19 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)72
2Gemini 3 Flash68
3DeepSeek V4 Pro67
4Claude Opus 4.662
5GPT-5.459
6GPT-5 Mini58
7DeepSeek R1 052855
8DeepSeek V4 Flash54
9Claude Sonnet 4.553
10Gemini 2.5 Flash53
11GPT-5.153
12Claude Sonnet 4.5 (Thinking)53
13Qwen 3 Max49
14Gemini 2.5 Pro47
15Qwen 3 32B32

Interactive version: theaggregate.ai/benchmark?slug=or-space · How It Works · Data refreshed daily, snapshot 2026-09-05.