FINESSE-Bench - Trading Derivatives: leaderboard

Metric: Accuracy (%) on the 544 Trading_derivatives numerical and short-answer tasks on options, put-call parity, Greeks, hedging, pricing and futures strategies of FINESSE-Bench; a GPT-5.2 judge marks each answer correct or incorrect against the reference; zero-shot, one fixed prompt per task type, temperature 0 where possible, reasoning configurations with medium effort where the model offers them; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 31 models tracked.

Top models

#ModelScore
1Kimi K2.586.58
2GPT-5.2 (Medium)85.11
3Qwen 3.5 Plus (2026-02-15)83.64
4Qwen 3.5 397B A17B81.25
5Qwen 3.5 122B A10B80.51
6GLM-580.51
7GLM-4.779.23
8Claude Sonnet 4.6 (Medium)78.49
9DeepSeek V3.277.39
10Qwen 3 235B A22B 2507 (Thinking)76.84
11DeepSeek R1 052876.1
12Qwen 3.5 35B A3B75.55
13Qwen 3.5 27B75.18
14GPT-5.4 (Medium)75.18
15Qwen 3.5 Flash (02-23)74.26

Interactive version: theaggregate.ai/benchmark?slug=finesse-bench-trading-derivatives · How It Works · Data refreshed daily, snapshot 2026-10-07.