Nous-Hermes-llama-2-7B: benchmark results

Provider: Nous Research. Released 2023-07-25. Access: Open.

Unified ELO 1342 ± 10, rank #2273 of 2656 rated models, from 85 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
MMLU-by-task - Business Ethics57Accuracy (%)81
MMLU-by-task - Professional Medicine54.41Accuracy (%)78
MMLU-by-task - Electrical Engineering49.66Accuracy (%)73.9
Hallucinations Leaderboard46.65Average Task Score (score)73.7
MMLU-by-task - TruthfulQA MC133.41Accuracy (%)72.8
MMLU-by-task - Machine Learning37.5Accuracy (%)71.9
MMLU-by-task - TruthfulQA MC249.01Accuracy (%)71.4
MMLU-by-task - Abstract Algebra32Accuracy (%)68.1
MMLU-by-task - College Computer Science44Accuracy (%)67.5
Open LLM Leaderboard - MuSR42.57Score66.1
MMLU-by-task - College Medicine45.66Accuracy (%)63.5
MMLU-by-task - Human Aging59.64Accuracy (%)63.5

Interactive version: theaggregate.ai/model?slug=nous-hermes-llama-2-7b · How It Works · Data refreshed daily, snapshot 2026-09-19.