Open FinLLM Reasoning - DM-Complong: leaderboard

Metric: Accuracy (%). Source: huggingface.co. 26 models tracked.

Top models

#ModelScore
1DeepSeek V342.33
2GPT-4o39.33
3GPT-4.539.33
4DeepSeek R138.67
5O136.67
6Llama 3.1 70B Instruct34.33
7Llama 3.3 70B Instruct32
8DeepSeek R1 Distill Llama 70B30.67
9Qwen 2.5 32B Instruct30
10Qwen 2.5 14B Instruct26.67
11DeepSeek R1 Distill Qwen 32B24.67
12DeepSeek R1 Distill Qwen 14B21
13QwQ-32B20
14Qwen 2.5 7B Instruct17.67
15DeepSeek R1 Distill Llama 8B15.67

Interactive version: theaggregate.ai/benchmark?slug=open-finllm-reasoning-dm-complong · How It Works · Data refreshed daily, snapshot 2026-09-05.