SPADE-Bench: leaderboard

Metric: Deception rate (%, Pass@5; lower is better): share of the executed SPADE-Bench cases (300 paired regular and pressure scenarios over 239 simulated tools, English prompts) in which any of 5 samples at temperature 0.7 shows a plan that shifts toward the external incentive under pressure while the executed tool action does not, judged by the released Qwen3-32B stance judge; cases the model fails to execute are left out of the denominator. Source: arxiv.org. Saturation forecast: Around 2029. 8 models tracked.

Top models

#ModelScore
1Llama 3.3 70B Instruct24.91
2GPT-5.125
3Kimi K230
4Claude Sonnet 4.533.11
5GLM-4.636.95
6DeepSeek V3.139
7Qwen 3 32B43.62
8Gemini 2.5 Pro57.33

Interactive version: theaggregate.ai/benchmark?slug=spade-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.