ArmoRM-Llama3-8B-v0.1 — benchmark results

Provider: Other. Released 2024-05-23. Access: Open.

Unified ELO 1372 ± 57, rank #1334 of 1776 rated models, from 27 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
RewardBench Reasoning97.35Accuracy (%)93.5
RewardBench Prior Sets (0.5 weight)74.29Score (%)93.3
RewardBench88.6Score (%)88.6
RewardBench Safety90.54Accuracy (%)86.1
RewardBench Precise IF41.88Score (%)83.2
RewardBench Chat96.93Accuracy (%)83
RewardBench Math66.12Score (%)78.8
RewardBench Chat Hard76.75Accuracy (%)75.6
RewardBench Ties66.29Score (%)66.3
Open Chinese LLM - TruthfulQA MC55.04Accuracy (%)61
RewardBench Focus76.57Score (%)59.2
StoryAlign49.7Average (self-reported)57.1

Interactive version: theaggregate.ai/model?slug=armorm-llama3-8b-v0-1 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.