CodeLlama-7B-Python-hf — benchmark results
Provider: Meta. Released 2023-08-24. Access: Open.
Unified ELO 1203 ± 15, rank #1689 of 1776 rated models, from 72 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MMLU-by-task - College Physics | 35.29 | Accuracy (%) | 88.7 |
| MMLU-by-task - High School Physics | 35.1 | Accuracy (%) | 81.9 |
| ToolBench - Google Sheets | 49.13 | Task Score | 79.1 |
| MMLU-by-task - High School Statistics | 46.3 | Accuracy (%) | 79 |
| ToolBench - Home Search | 83 | Task Score | 74.4 |
| ToolBench - The Cat API | 88 | Task Score | 74.4 |
| ToolBench - Trip Booking | 68.33 | Task Score | 72.1 |
| ToolBench Leaderboard | 52.15 | Average Task Score | 72.1 |
| MMLU-by-task - College Chemistry | 39 | Accuracy (%) | 70.2 |
| ToolBench - Tabletop | 22.86 | Task Score | 69.8 |
| ToolBench - VirtualHome | 21.33 | Task Score | 67.4 |
| ToolBench - WebShop Short | 1.58 | Task Score | 58.1 |
Interactive version: theaggregate.ai/model?slug=codellama-7b-python-hf · How the rankings work · Data refreshed daily, snapshot 2026-07-22.