FINESSE-Bench - CMT-like Level 2: leaderboard

Metric: Accuracy (%) on the 251 CMT-like Level 2 multiple-choice questions on technical analysis theory, chart patterns, indicators and trading-system testing of FINESSE-Bench; a GPT-5.2 judge marks each answer correct or incorrect against the reference; zero-shot, one fixed prompt per task type, temperature 0 where possible, reasoning configurations with medium effort where the model offers them; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 31 models tracked.

Top models

#ModelScore
1GLM-589.64
2Claude Sonnet 4.6 (Medium)89.24
3Qwen 3.5 Plus (2026-02-15)89.24
4Qwen 3.5 397B A17B88.84
5Kimi K2.588.05
6GLM-4.788.05
7Qwen 3.5 27B86.45
8Qwen 3.5 122B A10B86.06
9DeepSeek R1 052885.66
10GPT-5.4 (Medium)85.66
11GPT-5.2 (Medium)85.26
12Qwen 3 235B A22B 2507 (Thinking)84.46
13MiniMax-M2.184.46
14MiniMax-M2.584.06
15Qwen 3.5 35B A3B83.67

Interactive version: theaggregate.ai/benchmark?slug=finesse-bench-cmt-like-level-2 · How It Works · Data refreshed daily, snapshot 2026-10-07.