Open FinLLM Reasoning - XBRL-Math: leaderboard

Metric: Accuracy (%). Source: huggingface.co. 26 models tracked.

Top models

#ModelScore
1DeepSeek R186.67
2DeepSeek R1 Distill Llama 70B86.67
3QwQ-32B84.44
4DeepSeek R1 Distill Qwen 32B84.44
5DeepSeek R1 Distill Qwen 14B84.44
6Qwen2.5-Math-72B-Instruct83.33
7DeepSeek R1 Distill Llama 8B81.11
8DeepSeek V376.67
9O174.44
10GPT-4.574.44
11GPT-4o72.22
12Llama 3.3 70B Instruct70
13Qwen 2.5 72B Instruct67.78
14Qwen 2.5 32B Instruct65.56
15Llama 3.1 70B Instruct63.33

Interactive version: theaggregate.ai/benchmark?slug=open-finllm-reasoning-xbrl-math · How It Works · Data refreshed daily, snapshot 2026-09-05.