PortBench-QA - Max-Sharpe Allocation: leaderboard

Metric: Item score (%; long-only maximum-Sharpe weights for three or more assets with supplied means and covariances, 50 test questions per model). Source: arxiv.org. Saturation forecast: Estimated already saturated. 10 models tracked.

Top models

#ModelScore
1DeepSeek V4 Pro (Thinking)99.2
2Qwen 3.7 Max (Thinking)95.4
3DeepSeek V4 Flash (Thinking)93.2
4Qwen 3.6 Plus (Thinking)80.4
5GLM-5.1 (Thinking)42.1
6Kimi K2.6 (Thinking)28
7Qwen 3.6 35B A3B (Thinking)23

Interactive version: theaggregate.ai/benchmark?slug=portbench-qa-max-sharpe-allocation · How It Works · Data refreshed daily, snapshot 2026-09-26.