Open FinLLM Reasoning - FinQA: leaderboard

Metric: Accuracy (%). Source: huggingface.co. 26 models tracked.

Top models

#ModelScore
1Qwen 2.5 72B Instruct73.38
2DeepSeek V373.2
3Qwen 2.5 32B Instruct73.11
4GPT-4o72.49
5Qwen2.5-Math-72B-Instruct69.74
6GPT-4.568.94
7Llama 3.3 70B Instruct68.15
8Qwen 2.5 14B Instruct67.44
9DeepSeek R1 Distill Llama 70B66.73
10DeepSeek R1 Distill Qwen 32B65.48
11DeepSeek R165.13
12DeepSeek R1 Distill Qwen 14B63.27
13Llama 3.1 70B Instruct63.18
14QwQ-32B61.22
15Llama 3 70B Instruct58.92

Interactive version: theaggregate.ai/benchmark?slug=open-finllm-reasoning-finqa · How It Works · Data refreshed daily, snapshot 2026-09-05.