Open FinLLM Reasoning - XBRL-Math — leaderboard

Metric: Accuracy (%). Source: huggingface.co. 26 models tracked.

Top models

#ModelScore
1Fino1-14B87.78
2DeepSeek R186.67
3DeepSeek R1 Distill Llama 70B86.67
4QwQ-32B84.44
5DeepSeek R1 Distill Qwen 32B84.44
6DeepSeek R1 Distill Qwen 14B84.44
7Fino1-8B84.44
8Qwen2.5-Math-72B-Instruct83.33
9DeepSeek R1 Distill Llama 8B81.11
10DeepSeek V376.67
11O174.44
12GPT-4.574.44
13GPT-4o72.22
14Llama 3.3 70B Instruct70
15Qwen 2.5 72B Instruct67.78

Interactive version: theaggregate.ai/benchmark?slug=open-finllm-reasoning-xbrl-math · How the rankings work · Data refreshed daily, snapshot 2026-07-22.