HELM TORR - Fin Qa: leaderboard

Metric: Program Accuracy. Source: crfm.stanford.edu. 14 models tracked.

Top models

#ModelScore
1Gemini 1.5 Pro (002)47.06
2DeepSeek V346.29
3Claude 3.5 Sonnet (20241022)43.23
4Gemini 1.5 Flash (002)42.8
5GPT-4o (2024-11-20)39.63
6Qwen 2 72B Instruct37.91
7Llama 3.1 70B Instruct36.83
8Llama 3.1 405B Instruct35.77
9Claude 3.5 Haiku (20241022)35.43
10GPT-4o Mini (2024-07-18)33.51
11Mistral 7B Instruct (v0.3)19.43
12Llama 3.1 8B Instruct3.94

Interactive version: theaggregate.ai/benchmark?slug=helm-torr-fin-qa · How It Works · Data refreshed daily, snapshot 2026-09-08.