FinBen - QA — leaderboard
Metric: Normalized Score. Source: huggingface.co. 20 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4o | 78.22 |
| 2 | Llama 4 Scout Instruct | 74.22 |
| 3 | Llama 3.1 70B Instruct | 64.44 |
| 4 | plutus-8B-instruct | 64 |
| 5 | Qwen 2.5 32B Instruct | 60.44 |
| 6 | DeepSeek V3 | 50 |
| 7 | Qwen2.5-Omni-7B | 48.89 |
| 8 | finma-7B-full | 25.33 |
| 9 | Gemma 3 27B (IT) | 22.67 |
| 10 | Gemma 3 4B (IT) | 22.67 |
| 11 | Qwen2-Audio-7B-Instruct | 0 |
| 12 | SALMONN-7B | 0 |
Interactive version: theaggregate.ai/benchmark?slug=finben-qa · How the rankings work · Data refreshed daily, snapshot 2026-07-22.