EVMbench Incidents (Detect): leaderboard

Metric: Detect score (%): share of graded incidents (20 to 22 of the 22 real-world smart-contract incidents) whose single ground-truth vulnerability the agent's audit report identifies, graded by EVMbench's GPT-5 model-based detect grader; agents audit the deployed contracts without hints in Claude Code, Codex CLI or OpenCode (the Gemini 3.1 Pro run adds custom tools); higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 8 models tracked.

Top models

#ModelScoreOverall rank
1Claude Opus 4.665#60
2GPT-5.3 Codex (High)59.1#68 (GPT-5.3 Codex)
3Claude Sonnet 4.655#85
4GPT-5.3 Codex (xHigh)50#68 (GPT-5.3 Codex)
5GLM-542.9#137
6Gemini 3.1 Pro (Preview)30#54

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=evmbench-incidents-detect · How It Works · Data refreshed daily, snapshot 2026-10-11.