FinEval - Financial Industry Knowledge: leaderboard
Metric: Similarity score (%, zero-shot). Source: github.com. Saturation forecast: Estimated already saturated. 19 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4o | 61.3 |
| 2 | Gemini 1.5 Flash | 61.2 |
| 3 | GPT-4o Mini | 61.1 |
| 4 | Claude 3.5 Sonnet | 60.6 |
| 5 | Gemini 1.5 Pro | 60.5 |
| 6 | Qwen 2.5 72B Instruct | 54.4 |
| 7 | internlm2.5-20B-chat | 53.2 |
| 8 | GLM-4 9B Chat | 53.1 |
| 9 | internlm2-chat-20B | 50.3 |
| 10 | Baichuan2-13B-Chat | 50.2 |
| 11 | Yi 1.5 34B Chat | 49.6 |
| 12 | chatglm3-6B | 48.6 |
| 13 | Qwen 2.5 7B Instruct | 48.3 |
| 14 | Yi-1.5-9B Chat | 44.7 |
Interactive version: theaggregate.ai/benchmark?slug=fineval-financial-industry-knowledge · How It Works · Data refreshed daily, snapshot 2026-09-26.