BizFinBench — leaderboard

Financial benchmark evaluating LLMs across 9 business finance tasks: anomalous event attribution, financial computation, time reasoning, tool usage, QA, data description, emotion recognition, stock prediction, and named entity recognition.

Metric: Average Score. Source: github.com. Status: saturation imminent. 25 models tracked.

Top models

#ModelScore
1DeepSeek R173.05
2GPT-4o71.8
3DeepSeek V371.57
4Gemini 2.0 Flash69.75
5Qwen 3 32B68.26
6Qwen 2.5 72B Instruct67.7
7Qwen 3 14B67.05
8DeepSeek R1 Distill Qwen 32B66.29
9QwQ-32B65.77
10Claude 3.5 Sonnet65.59
11Llama 4 Scout61.17
12Qwen 3 4B59.94
13DeepSeek R1 Distill Qwen 14B59.49
14Qwen 2.5 7B Instruct56.35
15Llama 3.1 70B Instruct55.09

Interactive version: theaggregate.ai/benchmark?slug=bizfinbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.