FINESSE-Bench - CFA-like Level 1: leaderboard

Metric: Accuracy (%) on the 1069 CFA-like Level 1 multiple-choice questions on foundational finance (ethics, quantitative methods, economics, financial reporting, corporate finance, investments) of FINESSE-Bench; a GPT-5.2 judge marks each answer correct or incorrect against the reference; zero-shot, one fixed prompt per task type, temperature 0 where possible, reasoning configurations with medium effort where the model offers them; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 31 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.6 (Medium)89.79
2Kimi K2.589.24
3Qwen 3.5 Plus (2026-02-15)88.96
4GLM-588.59
5Qwen 3.5 122B A10B88.03
6GLM-4.787.65
7Qwen 3.5 397B A17B87.56
8GPT-5.2 (Medium)87.36
9Qwen 3 235B A22B 2507 (Thinking)87
10Qwen 3.5 Flash (02-23)86.62
11Qwen 3.5 35B A3B86.34
12Qwen 3.5 27B85.87
13DeepSeek R1 052885.87
14Claude 3.7 Sonnet (Thinking)84.19
15Qwen 3 32B84

Interactive version: theaggregate.ai/benchmark?slug=finesse-bench-cfa-like-level-1 · How It Works · Data refreshed daily, snapshot 2026-10-07.