Open FinLLM Reasoning: leaderboard

Metric: Accuracy (%). Source: huggingface.co. 26 models tracked.

Top models

#ModelScore
1DeepSeek V361.3
2GPT-4o61.01
3DeepSeek R160.87
4GPT-4.560.43
5DeepSeek R1 Distill Llama 70B59.27
6DeepSeek R1 Distill Qwen 32B57.4
7Qwen 2.5 32B Instruct56.17
8Llama 3.3 70B Instruct56.04
9O154.05
10Qwen 2.5 72B Instruct53.71
11DeepSeek R1 Distill Qwen 14B53.18
12QwQ-32B52.91
13Qwen 2.5 14B Instruct52.72
14Llama 3.1 70B Instruct52.21
15Qwen2.5-Math-72B-Instruct50.02

Interactive version: theaggregate.ai/benchmark?slug=open-finllm-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-05.