CodeLlama-34B-hf — benchmark results
Provider: Meta. Released 2023-08-24. Access: Open.
Unified ELO 1239 ± 17, rank #1643 of 1776 rated models, from 77 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ToolBench - Google Sheets | 64.29 | Task Score | 100 |
| ToolBench - VirtualHome | 24.65 | Task Score | 95.3 |
| ToolBench - Tabletop | 51.32 | Task Score | 93 |
| ToolBench - Trip Booking | 88.33 | Task Score | 93 |
| ToolBench Leaderboard | 62.94 | Average Task Score | 93 |
| ToolBench - Home Search | 88 | Task Score | 90.7 |
| ToolBench - Open Weather | 96.39 | Task Score | 88.4 |
| MMLU-by-task - Abstract Algebra | 34 | Accuracy (%) | 82.1 |
| ToolBench - WebShop Short | 5.53 | Task Score | 72.1 |
| MMLU-by-task - High School Physics | 33.11 | Accuracy (%) | 70.7 |
| MMLU-by-task - College Chemistry | 39 | Accuracy (%) | 70.2 |
| MMLU-by-task - College Mathematics | 34 | Accuracy (%) | 69.2 |
Interactive version: theaggregate.ai/model?slug=codellama-34b-hf · How the rankings work · Data refreshed daily, snapshot 2026-07-22.