FINESSE-Bench - Trading and Technical Analysis: leaderboard

Metric: Accuracy (%) over the trading and technical-analysis group of FINESSE-Bench: Trading_derivatives (544), Trading_TA (413) and CFTe-like Level 1 (781), the per-dataset accuracies weighted by question count; a GPT-5.2 judge marks each answer correct or incorrect against the reference; zero-shot, one fixed prompt per task type, temperature 0 where possible, reasoning configurations with medium effort where the model offers them; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 31 models tracked.

Top models

#ModelScore
1Kimi K2.584.98
2Qwen 3.5 Plus (2026-02-15)84.75
3GPT-5.2 (Medium)84.24
4Qwen 3.5 397B A17B83.31
5Qwen 3.5 122B A10B82.79
6Claude Sonnet 4.6 (Medium)82.45
7GLM-4.782.05
8GLM-581.3
9Qwen 3.5 27B80.38
10Qwen 3 235B A22B 2507 (Thinking)80.27
11GPT-5.4 (Medium)79.69
12Qwen 3.5 Flash (02-23)78.94
13Qwen 3.5 35B A3B78.83
14Claude 3.7 Sonnet (Thinking)78.54
15DeepSeek R1 052878.48

Interactive version: theaggregate.ai/benchmark?slug=finesse-bench-trading-and-technical-analysis · How It Works · Data refreshed daily, snapshot 2026-10-07.