DeepHermes-3-Llama-3-8B-Preview: benchmark results
Provider: Nous Research. Released 2025-02-12. Access: Open.
Unified ELO 1479 ± 65, rank #1290 of 2656 rated models, from 67 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EuroEval Portuguese NLU - MultiWikiQA PT | 73.82 | Reading comprehension Score (%) | 78.5 |
| EuroEval Spanish NLU - Sentiment Headlines ES | 45.89 | Sentiment classification Score (%) | 72.3 |
| EuroEval Portuguese NLU - HAREM | 47.29 | Named entity recognition Score (%) | 70.7 |
| EuroEval Portuguese NLU - SST-2 PT | 79.36 | Sentiment classification Score (%) | 62 |
| EuroEval Portuguese NLU | 51.92 | NLU Average Score (%) | 61.2 |
| EuroEval Dutch NLU - DBRD | 88.4 | Sentiment classification Score (%) | 58.4 |
| EuroEval Danish Knowledge | 62.72 | Knowledge Average Score (%) | 57.9 |
| EuroEval Italian NLU - MultiNERD IT | 64.6 | Named entity recognition Score (%) | 57.6 |
| EuroEval Danish Knowledge - Danish Citizen Tests | 62.72 | MCC (x100) | 54.1 |
| EuroEval Spanish NLU | 43.39 | NLU Average Score (%) | 53.9 |
| EuroEval Spanish NLU - CoNLL ES | 60.78 | Named entity recognition Score (%) | 51.7 |
| EuroEval Portuguese | 45.3 | Average Score (%) | 51.6 |
Interactive version: theaggregate.ai/model?slug=deephermes-3-llama-3-8b-preview · How It Works · Data refreshed daily, snapshot 2026-09-19.