PortBench-QA - Position Sizing: leaderboard

Metric: Item score (%; fixed-fractional sizing from a drawdown limit and 99% VaR, 50 test questions per model). Source: arxiv.org. Saturation forecast: Estimated already saturated. 10 models tracked.

Top models

#ModelScore
1Qwen 3.6 Plus (Thinking)96.8
2GLM-5.1 (Thinking)96.4
3DeepSeek V4 Pro (Thinking)96.3
4Qwen 3.6 35B A3B (Thinking)96.1
5Qwen 3.7 Max (Thinking)95.1
6DeepSeek V4 Flash (Thinking)94.5
7Kimi K2.6 (Thinking)49.3

Interactive version: theaggregate.ai/benchmark?slug=portbench-qa-position-sizing · How It Works · Data refreshed daily, snapshot 2026-09-26.