RPCBench - Error Group - Unverifiable Premise - Composite Premise Critique: leaderboard

Metric: Composite Score (0-100). Source: github.com. Saturation forecast: Around December 2026. 11 models tracked.

Top models

#ModelScore
1DeepSeek V4 Pro79.47
2GPT-5.578.31
3Qwen 3.5 122B A10B73.44
4DeepSeek V4 Flash72.81
5Qwen 3.5 Plus69.81
6Qwen 3.5 397B A17B69.68
7Qwen 3.5 35B A3B67.91
8Gemini 3.1 Pro (Preview)66.82
9Claude Sonnet 4.661.8
10Llama 3.1 70B Instruct21.32
11Llama 3.1 8B Instruct16.8

Interactive version: theaggregate.ai/benchmark?slug=rpcbench-error-group-unverifiable-premise-composite-premise-critique · How It Works · Data refreshed daily, snapshot 2026-09-24.