OLMo-7B-0724-Instruct-hf: benchmark results

Provider: Allen AI. Access: Open.

Unified ELO 1305 ± 25, rank #1310 of 1516 rated models, from 103 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AbstentionBench - underspecified context - QASPER - F145.48Abstention F1 (%)100
AbstentionBench - underspecified context - QASPER - Precision30.17Abstention Precision (%)100
AbstentionBench - stale - FreshQA - Recall89.36Abstention Recall (%)97.4
Open CoT - LSAT Reading Comprehension18.59CoT Gain (%)80.2
AbstentionBench - answer unknown - BB/Known unknowns - Recall100Abstention Recall (%)71.1
AbstentionBench - answer unknown - CoCoNot/Unsupported - Precision98.78Abstention Precision (%)63.2
AbstentionBench - false premise - FalseQA - Recall55.17Abstention Recall (%)63.2
AbstentionBench - false premise - QAQA - Recall48.77Abstention Recall (%)63.2
AbstentionBench - stale - FreshQA - F147.19Abstention F1 (%)63.2
AbstentionBench - false premise - FalseQA - F165.12Abstention F1 (%)57.9
AbstentionBench - underspecified context - ALCUNA - Recall70.73Abstention Recall (%)57.9
AbstentionBench - underspecified context - BB/Disambiguate - Precision52.38Abstention Precision (%)57.9

Interactive version: theaggregate.ai/model?slug=olmo-7b-0724-instruct-hf · How It Works · Data refreshed daily, snapshot 2026-09-24.