Hermes 4 405B — benchmark results
Nous Research's Hermes 4 open hybrid-reasoning fine-tune of Llama 3.1 405B. Provider: Nous Research. Released 2025-08-26. Access: Open.
Unified ELO 1648 ± 22, rank #373 of 1776 rated models, from 84 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AGC-Bench - proparalogy | 0.94 | Dataset z-score | 98.8 |
| AGC-Bench - science_analogies | 2.45 | Dataset z-score | 98.8 |
| AGC-Bench - story_quality | 1.31 | Dataset z-score | 98.8 |
| AGC-Bench - hypobench | 0.43 | Dataset z-score | 97.6 |
| AGC-Bench - Problem Solving | 0.73 | JRT z-score | 96.3 |
| AGC-Bench - ocw_connections | 1.11 | Dataset z-score | 96.3 |
| AGC-Bench - schnovel | 1.56 | Dataset z-score | 95.7 |
| UGI Leaderboard | 54.07 | UGI Score | 95.6 |
| AGC-Bench - metaphoric_analogies | 2.18 | Dataset z-score | 95.1 |
| Diplomacy: Betrayal Tendency | 91.7 | Betrayal Rate (%) | 93.8 |
| AGC-Bench - historical_analogy | 1.41 | Dataset z-score | 91.2 |
| AGC-Bench - riddlesense | 0.98 | Dataset z-score | 88.8 |
Interactive version: theaggregate.ai/model?slug=hermes-4-405b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.