DeepHermes-3-Llama-3-8B-Preview: benchmark results

Provider: Nous Research. Released 2025-02-12. Access: Open.

Unified ELO 1479 ± 65, rank #1290 of 2656 rated models, from 67 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EuroEval Portuguese NLU - MultiWikiQA PT73.82Reading comprehension Score (%)78.5
EuroEval Spanish NLU - Sentiment Headlines ES45.89Sentiment classification Score (%)72.3
EuroEval Portuguese NLU - HAREM47.29Named entity recognition Score (%)70.7
EuroEval Portuguese NLU - SST-2 PT79.36Sentiment classification Score (%)62
EuroEval Portuguese NLU51.92NLU Average Score (%)61.2
EuroEval Dutch NLU - DBRD88.4Sentiment classification Score (%)58.4
EuroEval Danish Knowledge62.72Knowledge Average Score (%)57.9
EuroEval Italian NLU - MultiNERD IT64.6Named entity recognition Score (%)57.6
EuroEval Danish Knowledge - Danish Citizen Tests62.72MCC (x100)54.1
EuroEval Spanish NLU43.39NLU Average Score (%)53.9
EuroEval Spanish NLU - CoNLL ES60.78Named entity recognition Score (%)51.7
EuroEval Portuguese45.3Average Score (%)51.6

Interactive version: theaggregate.ai/model?slug=deephermes-3-llama-3-8b-preview · How It Works · Data refreshed daily, snapshot 2026-09-19.