CodeLlama-13B-Python-hf — benchmark results
Provider: Meta. Released 2023-08-24. Access: Open.
Unified ELO 1191 ± 19, rank #1700 of 1776 rated models, from 72 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ToolBench - Tabletop | 41.53 | Task Score | 86 |
| ToolBench - The Cat API | 92.35 | Task Score | 86 |
| ToolBench - Home Search | 86 | Task Score | 83.7 |
| ToolBench - Google Sheets | 50.79 | Task Score | 81.4 |
| ToolBench - Open Weather | 93 | Task Score | 80.2 |
| ToolBench Leaderboard | 56.31 | Average Task Score | 79.1 |
| ToolBench - VirtualHome | 21.94 | Task Score | 76.7 |
| ToolBench - Trip Booking | 62.5 | Task Score | 62.8 |
| ToolBench - WebShop Short | 2.4 | Task Score | 62.8 |
| MMLU-by-task - Moral Scenarios | 27.15 | Accuracy (%) | 58.5 |
| MMLU-by-task - High School Statistics | 39.81 | Accuracy (%) | 54.6 |
| MMLU-by-task - High School Physics | 29.8 | Accuracy (%) | 46.5 |
Interactive version: theaggregate.ai/model?slug=codellama-13b-python-hf · How the rankings work · Data refreshed daily, snapshot 2026-07-22.