FINESSE-Bench - CFA-like Level 3: leaderboard

Metric: Accuracy (%) on the 318 CFA-like Level 3 multiple-choice questions on portfolio management, wealth planning, risk management and ethics cases of FINESSE-Bench; a GPT-5.2 judge marks each answer correct or incorrect against the reference; zero-shot, one fixed prompt per task type, temperature 0 where possible, reasoning configurations with medium effort where the model offers them; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 31 models tracked.

Top models

#ModelScore
1GLM-583.96
2Claude Sonnet 4.6 (Medium)82.39
3Qwen 3.5 Plus (2026-02-15)82.08
4Kimi K2.581.13
5GLM-4.780.19
6GPT-5.2 (Medium)80.19
7Claude 3.7 Sonnet (Thinking)78.93
8DeepSeek R1 052877.36
9Qwen 3.5 122B A10B76.73
10Qwen 3.5 27B76.42
11Qwen 3.5 35B A3B75.16
12Qwen 3 235B A22B 2507 (Thinking)74.84
13Qwen 3.5 Flash (02-23)74.53
14Qwen 3.5 397B A17B72.01
15Llama 4 Maverick71.7

Interactive version: theaggregate.ai/benchmark?slug=finesse-bench-cfa-like-level-3 · How It Works · Data refreshed daily, snapshot 2026-10-07.