FINESSE-Bench - CFTe-like Level 1: leaderboard

Metric: Accuracy (%) on the 781 CFTe-like Level 1 multiple-choice questions on basic technical analysis (chart types, trends, support and resistance, moving averages) of FINESSE-Bench; a GPT-5.2 judge marks each answer correct or incorrect against the reference; zero-shot, one fixed prompt per task type, temperature 0 where possible, reasoning configurations with medium effort where the model offers them; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 31 models tracked.

Top models

#ModelScore
1Qwen 3.5 122B A10B86.81
2Qwen 3.5 Plus (2026-02-15)86.81
3Qwen 3.5 397B A17B85.53
4Kimi K2.584.76
5Claude Sonnet 4.6 (Medium)84.64
6Qwen 3.5 27B84.38
7GLM-4.784.25
8Qwen 3 235B A22B 2507 (Thinking)83.99
9GPT-5.2 (Medium)83.99
10Qwen 3.5 Flash (02-23)83.1
11Claude 3.7 Sonnet (Thinking)82.59
12Llama 4 Maverick82.2
13GPT-5.4 (Medium)82.2
14Qwen 3.5 35B A3B81.18
15GLM-581.18

Interactive version: theaggregate.ai/benchmark?slug=finesse-bench-cfte-like-level-1 · How It Works · Data refreshed daily, snapshot 2026-10-07.