CodeLlama-7B-hf: benchmark results

Provider: Meta. Released 2023-08-24. Access: Open.

Unified ELO 1294 ± 1, rank #1353 of 1392 rated models, from 77 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
MMLU-by-task - College Physics38.24Accuracy (%)95.2
MMLU-by-task - High School Statistics47.22Accuracy (%)83.4
ToolBench - The Cat API92Task Score82.6
ToolBench - VirtualHome21.97Task Score81.4
MMLU-by-task - Abstract Algebra33Accuracy (%)75.6
MMLU-by-task - High School Mathematics29.63Accuracy (%)72.8
ToolBench - Google Sheets38.08Task Score67.4
ToolBench - Trip Booking63.33Task Score67.4
ToolBench - Open Weather86Task Score60.5
ToolBench - Home Search74Task Score58.1
EffiBench - NET2.95Normalized Execution Time57.3
ShaderMatch13.71Clone Match Rate (%)57.1

Interactive version: theaggregate.ai/model?slug=codellama-7b-hf · How It Works · Data refreshed daily, snapshot 2026-09-05.