Hermes 4 405B (Thinking) — benchmark results

Provider: Nous Research. Released 2025-08-26. Access: Open.

Unified ELO 1656 ± 33, rank #353 of 1776 rated models, from 38 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SpeechMap Compliance89.7% Requests Completed89.4
AA Omniscience - Humanities & Social Sciences32.8Accuracy (%)82.8
AA Omniscience - Health31Accuracy (%)82.4
AA MMLU-Pro82.9Accuracy (%)82.3
AA Omniscience - Business27Accuracy (%)80.7
AA Omniscience - Law24.3Accuracy (%)80.5
AA-Omniscience Accuracy30.17Accuracy (%)79.5
AA Omniscience - Science, Engineering & Mathematics35.1Accuracy (%)77
AA LiveCodeBench68.57Pass@1 (%)76.9
AA Omniscience - Software Engineering (SWE) - Julia28Accuracy (%)75.6
AA Omniscience - Software Engineering (SWE) - R26Accuracy (%)74.3
AA Omniscience - Software Engineering (SWE) - HTML44Accuracy (%)72.1

Interactive version: theaggregate.ai/model?slug=hermes-4-405b-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.