Open FinLLM Reasoning — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 26 models tracked.

Top models

#ModelScore
1DeepSeek V361.3
2GPT-4o61.01
3DeepSeek R160.87
4GPT-4.560.43
5Fino1-8B59.95
6DeepSeek R1 Distill Llama 70B59.27
7Fino1-14B58.1
8DeepSeek R1 Distill Qwen 32B57.4
9Qwen 2.5 32B Instruct56.17
10Llama 3.3 70B Instruct56.04
11O154.05
12Qwen 2.5 72B Instruct53.71
13DeepSeek R1 Distill Qwen 14B53.18
14QwQ-32B52.91
15Qwen 2.5 14B Instruct52.72

Interactive version: theaggregate.ai/benchmark?slug=open-finllm-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.