Backtrader-Bench (No Tools): leaderboard
Metric: Accuracy (%; mean of 10 runs on the curated set of 30 four-option multiple-choice questions whose answers come from executing Backtrader backtests of five trading strategies on AAPL daily data 2020-2024 (10 easy, 10 medium, 10 hard); answered from parametric knowledge through the Cursor SDK with every tool disabled; chance 25%). Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 73 |
| 2 | Claude Opus 4.7 | 73 |
| 3 | GPT-5.3 Codex | 69.3 |
| 4 | GPT-5.5 | 68 |
| 5 | Claude Sonnet 4.6 | 67.3 |
| 6 | Claude Opus 4.6 | 64.3 |
| 7 | Grok 4.3 | 59.3 |
| 8 | Kimi K2.5 | 55.7 |
| 9 | Gemini 2.5 Flash | 49 |
Interactive version: theaggregate.ai/benchmark?slug=backtrader-bench-no-tools · How It Works · Data refreshed daily, snapshot 2026-09-26.