Hermes 3 - Llama-3.1 70B — benchmark results
Provider: Nous Research. Released 2024-08-16. Access: Open.
Unified ELO 1236 ± 39, rank #1649 of 1776 rated models, from 116 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open LLM Leaderboard - MuSR | 23.43 | Score | 98.4 |
| AGC-Bench - unfun_corpus | 1.37 | Dataset z-score | 97.4 |
| AGC-Bench - creatset | 1.61 | Dataset z-score | 97 |
| Open LLM Leaderboard - BBH | 53.77 | Score | 96.3 |
| LiveBench Typos | 62 | Score | 93.1 |
| Open LLM Leaderboard - IFEval | 76.61 | Score | 92.3 |
| BenchBench | 84.51 | Aggregate Score (%) | 91.9 |
| LiveBench Table Reformat | 66 | Score | 91.7 |
| AGC-Bench - science_analogies | 1.5 | Dataset z-score | 91.2 |
| Open LLM Leaderboard - GPQA | 14.88 | Score | 90.8 |
| Open CoT - LogiQA | 7.99 | CoT Gain (%) | 88.9 |
| Open CoT - LSAT Logical Reasoning | 18.63 | CoT Gain (%) | 87.8 |
Interactive version: theaggregate.ai/model?slug=hermes-3-llama-3-1-70b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.