Hermes-3-Llama-3.1-8B — benchmark results

Nous Research's Hermes 3 fine-tune of Llama 3.1 8B, a steerable generalist assistant with function calling, JSON mode and roleplay focus (August 2024). Provider: Nous Research. Released 2024-07-28. Access: Open.

Unified ELO 1445 ± 21, rank #1021 of 1776 rated models, from 36 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LatamBoard - FLORES Bidirectional45.94Score (%)90.9
LatamBoard - Translation Score46.55Score (%)90.9
LatamBoard - Spanish WNLI76.06Score (%)83.3
LatamBoard - Spanish Escola70.85Score (%)81.8
LatamBoard - Spanish OpenBookQA37.2Score (%)81.8
LatamBoard - Spanish COPA83.4Score (%)78.8
LatamBoard - Spanish XNLI47.67Score (%)75.8
Open LLM Leaderboard - MuSR13.62Score74.9
Open LLM Leaderboard - IFEval61.7Score73.8
LatamBoard - OAB Exams49.02Score (%)69.7
LatamBoard - Portuguese Score86.22Score (%)69.7
LatamBoard - Spanish Score61.76Score (%)69.7

Interactive version: theaggregate.ai/model?slug=hermes-3-llama-3-1-8b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.