QuantCode-Bench: leaderboard

Metric: Judge pass rate (%), single turn: share of the tasks whose generated strategy compiles, backtests without runtime errors, places at least one trade and is judged by an LLM to implement the described strategy, on QuantCode-Bench (400 Backtrader trading-strategy tasks from Reddit, TradingView, StackExchange, GitHub and synthetic sources, multi-timeframe historical data); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 17 models tracked.

Top models

#ModelScore
1Claude Opus 4.675.8
2GPT-5.470.2
3Claude Sonnet 4.569.8
4GPT-5.2 Codex67.5
5GLM-565.4
6Claude Sonnet 4.665
7Kimi K2.564.8
8Gemini 3 Flash59.8
9Grok 4.1 Fast48.9
10DeepSeek V3.248.8
11Qwen 3 235B A22B48.2
12Qwen 3 Coder 30B A3B Instruct39.2
13Gemini 2.5 Flash31.2
14Qwen 3 14B25.2
15Qwen 3 8B18.5

Interactive version: theaggregate.ai/benchmark?slug=quantcode-bench · How It Works · Data refreshed daily, snapshot 2026-10-07.