Open FinLLM Reasoning - FinQA — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 26 models tracked.

Top models

#ModelScore
1Fino1-14B74.18
2Qwen 2.5 72B Instruct73.38
3DeepSeek V373.2
4Qwen 2.5 32B Instruct73.11
5Fino1-8B73.03
6GPT-4o72.49
7Qwen2.5-Math-72B-Instruct69.74
8GPT-4.568.94
9Llama 3.3 70B Instruct68.15
10Qwen 2.5 14B Instruct67.44
11DeepSeek R1 Distill Llama 70B66.73
12DeepSeek R1 Distill Qwen 32B65.48
13DeepSeek R165.13
14DeepSeek R1 Distill Qwen 14B63.27
15Llama 3.1 70B Instruct63.18

Interactive version: theaggregate.ai/benchmark?slug=open-finllm-reasoning-finqa · How the rankings work · Data refreshed daily, snapshot 2026-07-22.