RPCBench - Error Group - Underspecified Premise - Conditional Strategy Quality: leaderboard

Metric: Strategy Quality (0-100). Source: github.com. Saturation forecast: Around 2030. 11 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)67.33
2Qwen 3.5 397B A17B66.93
3Qwen 3.5 35B A3B66.34
4DeepSeek V4 Pro64.72
5Claude Sonnet 4.663.4
6Qwen 3.5 122B A10B63.19
7Qwen 3.5 Plus62.86
8DeepSeek V4 Flash62.85
9GPT-5.558.26
10Llama 3.1 70B Instruct58
11Llama 3.1 8B Instruct53.85

Interactive version: theaggregate.ai/benchmark?slug=rpcbench-error-group-underspecified-premise-conditional-strategy-quality · How It Works · Data refreshed daily, snapshot 2026-09-24.