ShiJianBench - Advisory Content Quality: leaderboard
Metric: Content score (0-100): 0.3 x accuracy + 0.7 x personalization of each compliant advisor message, rated 0-2 by a GPT-4o judge and mapped to 0-100, full horizon, over five matched runs of 27 simulated investors (risk profile x financial literacy x asset tier, weighted by 7,199 real users) in an LLM investor simulator replaying Chinese public-fund market traces, with GPT-4o judging advisor messages. Source: arxiv.org. Saturation forecast: Around December 2026. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek V4 Flash | 80.98 |
| 2 | Claude Sonnet 4.6 | 76.69 |
| 3 | GPT-5 | 68.17 |
Interactive version: theaggregate.ai/benchmark?slug=shijianbench-advisory-content-quality · How It Works · Data refreshed daily, snapshot 2026-09-29.