Open FinLLM Reasoning - DM-Simplong — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 26 models tracked.

Top models

#ModelScore
1GPT-4o60
2Qwen 2.5 72B Instruct59
3Qwen 2.5 14B Instruct59
4GPT-4.559
5Qwen 2.5 32B Instruct56
6O156
7Fino1-8B56
8DeepSeek R1 Distill Qwen 32B55
9Fino1-14B55
10Llama 3.3 70B Instruct54
11DeepSeek R153
12DeepSeek V353
13DeepSeek R1 Distill Llama 70B53
14Llama 3.1 70B Instruct48
15QwQ-32B46

Interactive version: theaggregate.ai/benchmark?slug=open-finllm-reasoning-dm-simplong · How the rankings work · Data refreshed daily, snapshot 2026-07-22.