Llama 3.1 Instruct Turbo 70B — benchmark results
Provider: Meta. Released 2024-07-23. Access: Open.
Unified ELO 1421 ± 72, rank #1138 of 1839 rated models, from 7 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Safety BBQ | 95.4 | BBQ accuracy (%) | 68.6 |
| HELM Safety XSTest | 94.5 | LM Evaluated Safety score (%) | 34.9 |
| HELM Safety SimpleSafetyTests | 92.5 | LM Evaluated Safety score (%) | 12.8 |
| HELM Safety | 84.5 | Mean score (self-reported) | 12.3 |
| HELM AIR-Bench | 42.5 | Refusal Rate (%) | 9.3 |
| HELM Safety Anthropic Red Team | 93.2 | LM Evaluated Safety score (%) | 8.1 |
| HELM Safety HarmBench | 46.9 | LM Evaluated Safety score (%) | 8.1 |
Interactive version: theaggregate.ai/model?slug=llama-3-1-instruct-turbo-70b · How It Works · Data refreshed daily, snapshot 2026-08-05.