Hermes-2-Pro-Llama-3-8B — benchmark results

Nous Research's Hermes 2 Pro on Llama 3 8B, tuned on cleaned OpenHermes 2.5 data for function calling and structured JSON output. Provider: Nous Research. Released 2024-04-30. Access: Open.

Unified ELO 1410 ± 30, rank #1175 of 1776 rated models, from 14 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SWE-Arena1002Elo Score80.6
Open PL LLM - Multiple Choice48.54Average Multiple-Choice Score (%)65
Open PL LLM - Generative56.82Average Generative Score (%)64.3
Open LLM Leaderboard - IFEval53.62Score64.1
Open PL LLM Leaderboard53.22Average Score (%)64
Open LLM Leaderboard - MuSR11.25Score58.5
Open LLM Leaderboard - BBH30.67Score54.1
Polish EQ-Bench54.57EQ-Bench Score52.5
Open PL LLM - RAG56.81Average RAG Score (%)51
Open LLM Leaderboard - GPQA5.7Score47.8
Open LLM Leaderboard - MATH Level 58.38Score42.2
Open LLM Leaderboard - MMLU-Pro22.8Score38.2

Interactive version: theaggregate.ai/model?slug=hermes-2-pro-llama-3-8b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.