Open FinLLM Reasoning - DM-Complong — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 26 models tracked.

Top models

#ModelScore
1DeepSeek V342.33
2GPT-4o39.33
3GPT-4.539.33
4DeepSeek R138.67
5O136.67
6Llama 3.1 70B Instruct34.33
7Llama 3.3 70B Instruct32
8DeepSeek R1 Distill Llama 70B30.67
9Qwen 2.5 32B Instruct30
10Fino1-14B27.33
11Qwen 2.5 14B Instruct26.67
12Fino1-8B26.33
13DeepSeek R1 Distill Qwen 32B24.67
14DeepSeek R1 Distill Qwen 14B21
15QwQ-32B20

Interactive version: theaggregate.ai/benchmark?slug=open-finllm-reasoning-dm-complong · How the rankings work · Data refreshed daily, snapshot 2026-07-22.