Llama 3 70B Chat: benchmark results
Provider: Meta. Released 2024-04-18. Access: Open.
Unified ELO 1566 ± 33, rank #766 of 2656 rated models, from 24 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SAD - Self-Recognition (Situating Prompt) | 73.69 | Score (%, higher is better) | 90 |
| SAD - Introspection (Situating Prompt) | 35.81 | Score (%, higher is better) | 85 |
| SAD - Predict Words | 40.56 | Self-prediction score (%, higher is better) | 81.8 |
| SAD - Introspection | 34.81 | Score (%, higher is better) | 80 |
| SAD - Self-Recognition | 65.69 | Score (%, higher is better) | 80 |
| SAD - Stages (Situating Prompt) | 48.44 | Score (%, higher is better) | 80 |
| SAD - Predict Words (Situating Prompt) | 38.56 | Self-prediction score (%, higher is better) | 77.3 |
| SAD - Influence | 65.31 | Score (%, higher is better) | 75 |
| SAD - Mini | 60.64 | Score (%, higher is better) | 75 |
| SAD - Stages | 46.91 | Score (%, higher is better) | 75 |
| SAD - Lite | 45.26 | Score (%, higher is better) | 70 |
| SAD - Lite (Situating Prompt) | 50.08 | Score (%, higher is better) | 70 |
Interactive version: theaggregate.ai/model?slug=llama-3-70b-chat · How It Works · Data refreshed daily, snapshot 2026-09-19.