RPCBench - Domain - MIND-small - Composite Premise Critique: leaderboard

Metric: Composite Score (0-100). Source: github.com. Saturation forecast: Around October 2027. 11 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.641.96
2DeepSeek V4 Pro41.49
3GPT-5.539.66
4Qwen 3.5 122B A10B39.01
5Qwen 3.5 Plus38.59
6Qwen 3.5 397B A17B38.24
7Gemini 3.1 Pro (Preview)38.07
8DeepSeek V4 Flash36.19
9Qwen 3.5 35B A3B35.13
10Llama 3.1 70B Instruct9.21
11Llama 3.1 8B Instruct6.33

Interactive version: theaggregate.ai/benchmark?slug=rpcbench-domain-mind-small-composite-premise-critique · How It Works · Data refreshed daily, snapshot 2026-09-24.