FINESSE-Bench - Trading TA: leaderboard

Metric: Accuracy (%) on the 413 Trading_TA numerical and short-answer applied technical-analysis tasks (patterns, momentum and mean reversion, entries and exits, backtesting) of FINESSE-Bench; a GPT-5.2 judge marks each answer correct or incorrect against the reference; zero-shot, one fixed prompt per task type, temperature 0 where possible, reasoning configurations with medium effort where the model offers them; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 31 models tracked.

Top models

#ModelScore
1GPT-5.2 (Medium)83.54
2Claude Sonnet 4.6 (Medium)83.54
3Kimi K2.583.29
4GLM-582.57
5Qwen 3.5 Plus (2026-02-15)82.32
6Qwen 3.5 397B A17B81.84
7GLM-4.781.6
8GPT-5.4 (Medium)80.87
9Qwen 3.5 27B79.66
10Qwen 3.5 35B A3B78.69
11Qwen 3.5 122B A10B78.21
12Claude 3.7 Sonnet (Thinking)77.97
13Qwen 3 235B A22B 2507 (Thinking)77.72
14Qwen 3.5 Flash (02-23)77.24
15DeepSeek R1 052876.51

Interactive version: theaggregate.ai/benchmark?slug=finesse-bench-trading-ta · How It Works · Data refreshed daily, snapshot 2026-10-07.