CodeLlama-34B-Python-hf — benchmark results
Provider: Meta. Released 2023-08-24. Access: Open.
Unified ELO 1226 ± 17, rank #1658 of 1776 rated models, from 73 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ToolBench - Home Search | 91 | Task Score | 95.3 |
| ToolBench - Trip Booking | 85.83 | Task Score | 90.7 |
| ToolBench - Google Sheets | 55.87 | Task Score | 88.4 |
| ToolBench Leaderboard | 59.16 | Average Task Score | 86 |
| ToolBench - Tabletop | 33.33 | Task Score | 79.1 |
| ToolBench - Open Weather | 91.11 | Task Score | 76.7 |
| ToolBench - The Cat API | 88.42 | Task Score | 76.7 |
| ToolBench - WebShop Short | 6.47 | Task Score | 74.4 |
| MMLU-by-task - Machine Learning | 37.5 | Accuracy (%) | 71.9 |
| MMLU-by-task - Formal Logic | 34.92 | Accuracy (%) | 69.2 |
| ToolBench - VirtualHome | 21.24 | Task Score | 65.1 |
| MMLU-by-task - Elementary Mathematics | 31.48 | Accuracy (%) | 63.7 |
Interactive version: theaggregate.ai/model?slug=codellama-34b-python-hf · How the rankings work · Data refreshed daily, snapshot 2026-07-22.