SPADE-Bench (Chinese): leaderboard

Metric: Deception rate (%, Pass@5; lower is better): share of the executed SPADE-Bench cases (300 paired regular and pressure scenarios over 239 simulated tools, Chinese prompts) in which any of 5 samples at temperature 0.7 shows a plan that shifts toward the external incentive under pressure while the executed tool action does not, judged by the released Qwen3-32B stance judge; cases the model fails to execute are left out of the denominator. Source: arxiv.org. Saturation forecast: Around August 2028. 4 models tracked.

Top models

#ModelScore
1Llama 3.3 70B Instruct18.32
2GLM-4.635.12
3Qwen 3 32B42.47
4Gemini 2.5 Pro46.49

Interactive version: theaggregate.ai/benchmark?slug=spade-bench-chinese · How It Works · Data refreshed daily, snapshot 2026-09-29.