CodeLlama-7B-hf — benchmark results
Provider: Meta. Released 2023-08-24. Access: Open.
Unified ELO 1207 ± 15, rank #1682 of 1776 rated models, from 77 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MMLU-by-task - College Physics | 38.24 | Accuracy (%) | 95.2 |
| MMLU-by-task - High School Statistics | 47.22 | Accuracy (%) | 83.4 |
| ToolBench - The Cat API | 92 | Task Score | 82.6 |
| ToolBench - VirtualHome | 21.97 | Task Score | 81.4 |
| MMLU-by-task - Abstract Algebra | 33 | Accuracy (%) | 75.6 |
| MMLU-by-task - High School Mathematics | 29.63 | Accuracy (%) | 72.8 |
| ToolBench - Google Sheets | 38.08 | Task Score | 67.4 |
| ToolBench - Trip Booking | 63.33 | Task Score | 67.4 |
| ToolBench - Open Weather | 86 | Task Score | 60.5 |
| ToolBench - Home Search | 74 | Task Score | 58.1 |
| EffiBench - NET | 2.95 | Normalized Execution Time | 57.3 |
| ShaderMatch | 13.71 | Clone Match Rate (%) | 57.1 |
Interactive version: theaggregate.ai/model?slug=codellama-7b-hf · How the rankings work · Data refreshed daily, snapshot 2026-07-22.