Rigged Backtest Audit - All-Three Specificity: leaderboard
Metric: All-three specificity (%; share of the 48 flawed backtests for which the audit names the right flaw class, cites the planted lines and proposes a class-specific fix; clean-aware prompt on the code surface of a 96-item paired set, 48 flawed backtests over eight flaw classes and 48 matched clean controls, deterministic scorer). Source: arxiv.org. Saturation forecast: Estimated already saturated. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4.1 Mini | 97.9 |
| 2 | DeepSeek V4 Flash | 87.5 |
| 3 | Gemini 2.5 Flash Lite | 85.4 |
| 4 | GPT-4o Mini | 75 |
Interactive version: theaggregate.ai/benchmark?slug=rigged-backtest-audit-all-three-specificity · How It Works · Data refreshed daily, snapshot 2026-09-26.