Hermes-3-Llama-3.1-8B — benchmark results
Nous Research's Hermes 3 fine-tune of Llama 3.1 8B, a steerable generalist assistant with function calling, JSON mode and roleplay focus (August 2024). Provider: Nous Research. Released 2024-07-28. Access: Open.
Unified ELO 1445 ± 21, rank #1021 of 1776 rated models, from 36 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LatamBoard - FLORES Bidirectional | 45.94 | Score (%) | 90.9 |
| LatamBoard - Translation Score | 46.55 | Score (%) | 90.9 |
| LatamBoard - Spanish WNLI | 76.06 | Score (%) | 83.3 |
| LatamBoard - Spanish Escola | 70.85 | Score (%) | 81.8 |
| LatamBoard - Spanish OpenBookQA | 37.2 | Score (%) | 81.8 |
| LatamBoard - Spanish COPA | 83.4 | Score (%) | 78.8 |
| LatamBoard - Spanish XNLI | 47.67 | Score (%) | 75.8 |
| Open LLM Leaderboard - MuSR | 13.62 | Score | 74.9 |
| Open LLM Leaderboard - IFEval | 61.7 | Score | 73.8 |
| LatamBoard - OAB Exams | 49.02 | Score (%) | 69.7 |
| LatamBoard - Portuguese Score | 86.22 | Score (%) | 69.7 |
| LatamBoard - Spanish Score | 61.76 | Score (%) | 69.7 |
Interactive version: theaggregate.ai/model?slug=hermes-3-llama-3-1-8b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.