HELM TORR - Wikitq - Robustness: leaderboard

Metric: Robustness (0-100). Source: crfm.stanford.edu. 14 models tracked.

Top models

#ModelScore
1Claude 3.5 Sonnet (20241022)80.19
2DeepSeek V378.21
3Llama 3.1 405B Instruct76.09
4Llama 3.1 70B Instruct75.87
5Qwen 2 72B Instruct75.34
6GPT-4o (2024-11-20)75.33
7GPT-4o Mini (2024-07-18)73.3
8Gemini 1.5 Flash (002)72.33
9Gemini 1.5 Pro (002)69.01
10Llama 3.1 8B Instruct66.83
11Claude 3.5 Haiku (20241022)60.78
12Mistral 7B Instruct (v0.3)39.85

Interactive version: theaggregate.ai/benchmark?slug=helm-torr-wikitq-robustness · How It Works · Data refreshed daily, snapshot 2026-09-19.