QuantCode-Bench (Agentic): leaderboard
Metric: Judge pass rate (%) within up to 10 attempts, each failed attempt returned to the model with its error type and system message, on QuantCode-Bench (400 Backtrader trading-strategy tasks from Reddit, TradingView, StackExchange, GitHub and synthetic sources, multi-timeframe historical data); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 17 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.6 | 97.5 |
| 2 | Claude Sonnet 4.6 | 96 |
| 3 | GPT-5.4 | 95 |
| 4 | Kimi K2.5 | 93.5 |
| 5 | Claude Sonnet 4.5 | 93 |
| 6 | Gemini 3 Flash | 91.8 |
| 7 | GLM-5 | 90.8 |
| 8 | GPT-5.2 Codex | 89.8 |
| 9 | Qwen 3 235B A22B | 87.2 |
| 10 | Grok 4.1 Fast | 84.5 |
| 11 | DeepSeek V3.2 | 83.8 |
| 12 | Qwen 3 Coder 30B A3B Instruct | 68 |
| 13 | Qwen 3 14B | 62.7 |
| 14 | Gemini 2.5 Flash | 62.5 |
| 15 | Qwen 3 8B | 47.6 |
Interactive version: theaggregate.ai/benchmark?slug=quantcode-bench-agentic · How It Works · Data refreshed daily, snapshot 2026-10-07.