BacktestBench - SQL Execution Accuracy: leaderboard
Metric: Execution accuracy (%) of the intermediate SQL: the share of generated queries whose result set matches the ground-truth result set ignoring order, synthetic BacktestBench test split (the expert-crafted subset excluded), each model driving the AutoBacktest pipeline of the authors (Summarizer, SQL Retriever and Python Coder agents) over a PostgreSQL market database, temperature 0.6; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 23 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Pro | 98.7 |
| 2 | Seed 1.8 | 97.45 |
| 3 | Qwen 3 235B A22B 2507 (Thinking) | 97.11 |
| 4 | GPT-OSS-120B | 96.58 |
| 5 | Kimi K2 (Thinking) | 96.06 |
| 6 | MiniMax-M2.1 | 95.02 |
| 7 | GLM-4.7 | 94.86 |
| 8 | Qwen 3 Next 80B A3B (Thinking) | 94.79 |
| 9 | Qwen 3 Max | 94.73 |
| 10 | Qwen 3 Coder Plus | 94.03 |
| 11 | DeepSeek V3.2 | 92.86 |
| 12 | Qwen 3 30B A3B 2507 (Thinking) | 92.86 |
| 13 | Qwen 3 32B | 90.72 |
| 14 | GPT-OSS-20B | 90.56 |
| 15 | Qwen 3 14B | 87.59 |
Interactive version: theaggregate.ai/benchmark?slug=backtestbench-sql-execution-accuracy · How It Works · Data refreshed daily, snapshot 2026-10-07.