Open FinLLM Reasoning - DM-Simplong: leaderboard

Metric: Accuracy (%). Source: huggingface.co. 26 models tracked.

Top models

#ModelScore
1GPT-4o60
2Qwen 2.5 72B Instruct59
3Qwen 2.5 14B Instruct59
4GPT-4.559
5Qwen 2.5 32B Instruct56
6O156
7DeepSeek R1 Distill Qwen 32B55
8Llama 3.3 70B Instruct54
9DeepSeek R153
10DeepSeek V353
11DeepSeek R1 Distill Llama 70B53
12Llama 3.1 70B Instruct48
13QwQ-32B46
14DeepSeek R1 Distill Qwen 14B44
15Qwen2.5-Math-72B-Instruct42

Interactive version: theaggregate.ai/benchmark?slug=open-finllm-reasoning-dm-simplong · How It Works · Data refreshed daily, snapshot 2026-09-05.