RPCBench - Error Group - Unverifiable Premise - Conditional Localization: leaderboard

Metric: Localization Accuracy (%). Source: github.com. Saturation forecast: Estimated already saturated. 11 models tracked.

Top models

#ModelScore
1GPT-5.593.64
2DeepSeek V4 Pro90.2
3Qwen 3.5 Plus89.98
4DeepSeek V4 Flash89.81
5Qwen 3.5 122B A10B89.62
6Qwen 3.5 397B A17B89.49
7Claude Sonnet 4.688.38
8Qwen 3.5 35B A3B87.95
9Gemini 3.1 Pro (Preview)83.86
10Llama 3.1 70B Instruct51.42
11Llama 3.1 8B Instruct50.69

Interactive version: theaggregate.ai/benchmark?slug=rpcbench-error-group-unverifiable-premise-conditional-localization · How It Works · Data refreshed daily, snapshot 2026-09-24.