FINESSE-Bench - CFA-like Level 2: leaderboard

Metric: Accuracy (%) on the 293 CFA-like Level 2 multiple-choice item sets on applied valuation, financial statement analysis, fixed income and derivatives of FINESSE-Bench; a GPT-5.2 judge marks each answer correct or incorrect against the reference; zero-shot, one fixed prompt per task type, temperature 0 where possible, reasoning configurations with medium effort where the model offers them; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 31 models tracked.

Top models

#ModelScore
1Kimi K2.591.81
2Claude Sonnet 4.6 (Medium)91.81
3GLM-4.791.47
4Qwen 3.5 Plus (2026-02-15)91.47
5GLM-589.76
6Qwen 3.5 27B89.08
7Qwen 3.5 122B A10B88.74
8GPT-5.2 (Medium)88.74
9Qwen 3.5 Flash (02-23)88.74
10Qwen 3.5 35B A3B88.4
11Qwen 3 235B A22B 2507 (Thinking)86.35
12DeepSeek R1 052885.32
13Claude 3.7 Sonnet (Thinking)83.96
14Qwen 3.5 397B A17B82.25
15MiniMax-M2.581.57

Interactive version: theaggregate.ai/benchmark?slug=finesse-bench-cfa-like-level-2 · How It Works · Data refreshed daily, snapshot 2026-10-07.