RPCBench - Error Group - Inconsistent Premise - Conditional Strategy Quality: leaderboard

Metric: Strategy Quality (0-100). Source: github.com. Saturation forecast: Estimated already saturated. 11 models tracked.

Top models

#ModelScore
1Qwen 3.5 397B A17B91.35
2Qwen 3.5 Plus91.23
3DeepSeek V4 Pro90.66
4Qwen 3.5 122B A10B90.16
5DeepSeek V4 Flash89.76
6Qwen 3.5 35B A3B89.34
7Claude Sonnet 4.688.61
8Gemini 3.1 Pro (Preview)88.1
9GPT-5.586.12
10Llama 3.1 70B Instruct62.93
11Llama 3.1 8B Instruct54.91

Interactive version: theaggregate.ai/benchmark?slug=rpcbench-error-group-inconsistent-premise-conditional-strategy-quality · How It Works · Data refreshed daily, snapshot 2026-09-24.