Backtrader-Bench - Mined Questions (No Tools): leaderboard
Metric: Accuracy (%; one run on the 38 four-option questions mined by a GPT-5.5 generator with executable verification and kept only when a tool-free GPT-5.4 solver failed and a tool-using solver succeeded; answered through the Cursor SDK with every tool disabled; chance 25%). Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 60.5 |
| 2 | GPT-5.5 | 60.5 |
| 3 | Claude Opus 4.6 | 57.9 |
| 4 | Claude Opus 4.7 | 52.6 |
| 5 | Claude Sonnet 4.6 | 39.5 |
| 6 | GPT-5.3 Codex | 34.2 |
| 7 | Gemini 2.5 Flash | 29 |
| 8 | Grok 4.3 | 21.1 |
| 9 | Kimi K2.5 | 15.8 |
Interactive version: theaggregate.ai/benchmark?slug=backtrader-bench-mined-questions-no-tools · How It Works · Data refreshed daily, snapshot 2026-09-26.