RPCBench - Domain - MIND-small - Evidence Distortion: leaderboard

Metric: Evidence Distortion Rate (%). Source: github.com. Saturation forecast: Around June 2028. 11 models tracked.

Top models

#ModelScore
1GPT-5.519.63
2Llama 3.1 70B Instruct23.47
3Gemini 3.1 Pro (Preview)23.61
4Qwen 3.5 397B A17B24.75
5Qwen 3.5 122B A10B25.75
6Qwen 3.5 35B A3B25.89
7Llama 3.1 8B Instruct27.88
8Qwen 3.5 Plus27.88
9DeepSeek V4 Flash28.88
10Claude Sonnet 4.630.3
11DeepSeek V4 Pro30.3

Interactive version: theaggregate.ai/benchmark?slug=rpcbench-domain-mind-small-evidence-distortion · How It Works · Data refreshed daily, snapshot 2026-09-24.