SPADE-Bench (Chinese): leaderboard
Metric: Deception rate (%, Pass@5; lower is better): share of the executed SPADE-Bench cases (300 paired regular and pressure scenarios over 239 simulated tools, Chinese prompts) in which any of 5 samples at temperature 0.7 shows a plan that shifts toward the external incentive under pressure while the executed tool action does not, judged by the released Qwen3-32B stance judge; cases the model fails to execute are left out of the denominator. Source: arxiv.org. Saturation forecast: Around August 2028. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Llama 3.3 70B Instruct | 18.32 |
| 2 | GLM-4.6 | 35.12 |
| 3 | Qwen 3 32B | 42.47 |
| 4 | Gemini 2.5 Pro | 46.49 |
Interactive version: theaggregate.ai/benchmark?slug=spade-bench-chinese · How It Works · Data refreshed daily, snapshot 2026-09-29.