HELM TORR - Fin Qa - Robustness: leaderboard

Metric: Robustness (0-100). Source: crfm.stanford.edu. 14 models tracked.

Top models

#ModelScore
1Llama 3.1 8B Instruct63
2Gemini 1.5 Pro (002)59
3Gemini 1.5 Flash (002)57
4Claude 3.5 Sonnet (20241022)56
5GPT-4o (2024-11-20)55
6Mistral 7B Instruct (v0.3)51
7DeepSeek V351
8Qwen 2 72B Instruct51
9Claude 3.5 Haiku (20241022)50
10Llama 3.1 70B Instruct49
11Llama 3.1 405B Instruct49
12GPT-4o Mini (2024-07-18)47

Interactive version: theaggregate.ai/benchmark?slug=helm-torr-fin-qa-robustness · How It Works · Data refreshed daily, snapshot 2026-09-19.