SCDBench: leaderboard
Metric: Semantically consistent public functions (%; all replayed test cases match return data, revert status, logs and storage, of 14,553 functions in 600 contracts, zero-shot). Source: arxiv.org. Saturation forecast: Around 2029. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.7 (High) | 22.48 |
| 2 | GLM-5 (Thinking) | 9.91 |
| 3 | GLM-5 (Non-reasoning) | 5.77 |
Interactive version: theaggregate.ai/benchmark?slug=scdbench · How It Works · Data refreshed daily, snapshot 2026-09-26.