PortBench-QA - Rebalancing: leaderboard

Metric: Item score (%; rebalance-or-hold decision with the corrective trade's asset and direction, partial credit, 50 test questions per model). Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.

Top models

#ModelScore
1GLM-5.1 (Thinking)88.2
2Qwen 3.7 Max (Thinking)72.4
3Kimi K2.6 (Thinking)68.4
4DeepSeek V4 Pro (Thinking)65.2
5DeepSeek V4 Flash (Thinking)65.2
6Qwen 3.6 Plus (Thinking)64
7Qwen 3.6 35B A3B (Thinking)56.4

Interactive version: theaggregate.ai/benchmark?slug=portbench-qa-rebalancing · How It Works · Data refreshed daily, snapshot 2026-09-26.