CodeLlama-34B-Instruct-hf — benchmark results
Provider: Meta. Released 2023-08-24. Access: Open.
Unified ELO 1303 ± 21, rank #1527 of 1776 rated models, from 89 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MMLU-by-task - College Mathematics | 41 | Accuracy (%) | 97.3 |
| ToolBench - Trip Booking | 89.17 | Task Score | 95.3 |
| ToolBench Leaderboard | 64.79 | Average Task Score | 95.3 |
| ToolBench - Google Sheets | 61.11 | Task Score | 93 |
| ToolBench - Home Search | 90 | Task Score | 93 |
| ToolBench - VirtualHome | 24.34 | Task Score | 93 |
| MMLU-by-task - Abstract Algebra | 36 | Accuracy (%) | 91.7 |
| ToolBench - Tabletop | 47.46 | Task Score | 90.7 |
| HumanLikeness - Sound-1 | 78.53 | Humanlike Score (%) | 89.5 |
| ToolBench - WebShop Short | 25.99 | Task Score | 88.4 |
| HumanLikeness - Discourse-1 | 79.12 | Humanlike Score (%) | 84.2 |
| HumanLikeness - Sound-2 | 63.97 | Humanlike Score (%) | 84.2 |
Interactive version: theaggregate.ai/model?slug=codellama-34b-instruct-hf · How the rankings work · Data refreshed daily, snapshot 2026-07-22.