Hermes 4 70B (Thinking) — benchmark results
Provider: Nous Research. Released 2025-08-26. Access: Open.
Unified ELO 1619 ± 35, rank #437 of 1776 rated models, from 37 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SpeechMap Compliance | 89.1 | % Requests Completed | 88.6 |
| AA MMLU-Pro | 81.05 | Accuracy (%) | 73.3 |
| AA LiveCodeBench | 65.29 | Pass@1 (%) | 72.5 |
| AA Omniscience - Humanities & Social Sciences | 26.2 | Accuracy (%) | 71 |
| AA Omniscience - Health | 24.1 | Accuracy (%) | 69.5 |
| AA Omniscience - Science, Engineering & Mathematics | 30.3 | Accuracy (%) | 66.2 |
| AA Omniscience - Software Engineering (SWE) - Julia | 20 | Accuracy (%) | 65.4 |
| AA Omniscience - Business | 20.4 | Accuracy (%) | 65.3 |
| AA Omniscience - Software Engineering (SWE) - HTML | 38 | Accuracy (%) | 65.3 |
| AA Omniscience - Law | 14.9 | Accuracy (%) | 64 |
| AA Omniscience - Software Engineering (SWE) - R | 18 | Accuracy (%) | 63.8 |
| AA AIME 2025 | 68.67 | Accuracy (%) | 63.5 |
Interactive version: theaggregate.ai/model?slug=hermes-4-70b-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.