DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning) — benchmark results
Provider: Nous Research. Released 2025-01-15. Access: Open.
Unified ELO 1228 ± 84, rank #1714 of 1841 rated models, from 7 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA Humanity's Last Exam | 4.26 | Accuracy (%) | 17.7 |
| Epoch AI - Scicode | 9.14 | Score | 8.8 |
| AA LiveCodeBench | 8.47 | Pass@1 (%) | 6.1 |
| AA MMLU-Pro | 36.53 | Accuracy (%) | 5.2 |
| Artificial Analysis Intelligence Index | 2.29 | Intelligence Index | 4 |
| AA GPQA Diamond | 26.97 | Accuracy (%) | 3.8 |
| AA MATH-500 | 21.8 | Accuracy (%) | 2.9 |
Interactive version: theaggregate.ai/model?slug=deephermes-3-llama-3-1-8b-preview-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-07-25.