Hermes 4 70B (Thinking) — benchmark results

Provider: Nous Research. Released 2025-08-26. Access: Open.

Unified ELO 1619 ± 35, rank #437 of 1776 rated models, from 37 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SpeechMap Compliance89.1% Requests Completed88.6
AA MMLU-Pro81.05Accuracy (%)73.3
AA LiveCodeBench65.29Pass@1 (%)72.5
AA Omniscience - Humanities & Social Sciences26.2Accuracy (%)71
AA Omniscience - Health24.1Accuracy (%)69.5
AA Omniscience - Science, Engineering & Mathematics30.3Accuracy (%)66.2
AA Omniscience - Software Engineering (SWE) - Julia20Accuracy (%)65.4
AA Omniscience - Business20.4Accuracy (%)65.3
AA Omniscience - Software Engineering (SWE) - HTML38Accuracy (%)65.3
AA Omniscience - Law14.9Accuracy (%)64
AA Omniscience - Software Engineering (SWE) - R18Accuracy (%)63.8
AA AIME 202568.67Accuracy (%)63.5

Interactive version: theaggregate.ai/model?slug=hermes-4-70b-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.