RPCBench - Conditional Localization: leaderboard

Metric: Localization Accuracy (%). Source: arxiv.org. Saturation forecast: Estimated already saturated. 11 models tracked.

Top models

#ModelScore
1GPT-5.591.72
2Qwen 3.5 Plus91.32
3Claude Sonnet 4.691.22
4Qwen 3.5 397B A17B91.11
5Qwen 3.5 122B A10B90.25
6DeepSeek V4 Pro88.59
7Gemini 3.1 Pro (Preview)88.53
8DeepSeek V4 Flash88.44
9Qwen 3.5 35B A3B88.4
10Llama 3.1 70B Instruct60.74
11Llama 3.1 8B Instruct55.33

Interactive version: theaggregate.ai/benchmark?slug=rpcbench-conditional-localization · How It Works · Data refreshed daily, snapshot 2026-09-24.