QuantCode-Bench (Agentic): leaderboard

Metric: Judge pass rate (%) within up to 10 attempts, each failed attempt returned to the model with its error type and system message, on QuantCode-Bench (400 Backtrader trading-strategy tasks from Reddit, TradingView, StackExchange, GitHub and synthetic sources, multi-timeframe historical data); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 17 models tracked.

Top models

#ModelScore
1Claude Opus 4.697.5
2Claude Sonnet 4.696
3GPT-5.495
4Kimi K2.593.5
5Claude Sonnet 4.593
6Gemini 3 Flash91.8
7GLM-590.8
8GPT-5.2 Codex89.8
9Qwen 3 235B A22B87.2
10Grok 4.1 Fast84.5
11DeepSeek V3.283.8
12Qwen 3 Coder 30B A3B Instruct68
13Qwen 3 14B62.7
14Gemini 2.5 Flash62.5
15Qwen 3 8B47.6

Interactive version: theaggregate.ai/benchmark?slug=quantcode-bench-agentic · How It Works · Data refreshed daily, snapshot 2026-10-07.