Rigged Backtest Audit - All-Three Specificity: leaderboard

Metric: All-three specificity (%; share of the 48 flawed backtests for which the audit names the right flaw class, cites the planted lines and proposes a class-specific fix; clean-aware prompt on the code surface of a 96-item paired set, 48 flawed backtests over eight flaw classes and 48 matched clean controls, deterministic scorer). Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.

Top models

#ModelScore
1GPT-4.1 Mini97.9
2DeepSeek V4 Flash87.5
3Gemini 2.5 Flash Lite85.4
4GPT-4o Mini75

Interactive version: theaggregate.ai/benchmark?slug=rigged-backtest-audit-all-three-specificity · How It Works · Data refreshed daily, snapshot 2026-09-26.