SPADE-Bench: leaderboard
Metric: Deception rate (%, Pass@5; lower is better): share of the executed SPADE-Bench cases (300 paired regular and pressure scenarios over 239 simulated tools, English prompts) in which any of 5 samples at temperature 0.7 shows a plan that shifts toward the external incentive under pressure while the executed tool action does not, judged by the released Qwen3-32B stance judge; cases the model fails to execute are left out of the denominator. Source: arxiv.org. Saturation forecast: Around 2029. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Llama 3.3 70B Instruct | 24.91 |
| 2 | GPT-5.1 | 25 |
| 3 | Kimi K2 | 30 |
| 4 | Claude Sonnet 4.5 | 33.11 |
| 5 | GLM-4.6 | 36.95 |
| 6 | DeepSeek V3.1 | 39 |
| 7 | Qwen 3 32B | 43.62 |
| 8 | Gemini 2.5 Pro | 57.33 |
Interactive version: theaggregate.ai/benchmark?slug=spade-bench · How It Works · Data refreshed daily, snapshot 2026-09-29.