CodeLlama-34B-Instruct-hf — benchmark results

Provider: Meta. Released 2023-08-24. Access: Open.

Unified ELO 1303 ± 21, rank #1527 of 1776 rated models, from 89 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
MMLU-by-task - College Mathematics41Accuracy (%)97.3
ToolBench - Trip Booking89.17Task Score95.3
ToolBench Leaderboard64.79Average Task Score95.3
ToolBench - Google Sheets61.11Task Score93
ToolBench - Home Search90Task Score93
ToolBench - VirtualHome24.34Task Score93
MMLU-by-task - Abstract Algebra36Accuracy (%)91.7
ToolBench - Tabletop47.46Task Score90.7
HumanLikeness - Sound-178.53Humanlike Score (%)89.5
ToolBench - WebShop Short25.99Task Score88.4
HumanLikeness - Discourse-179.12Humanlike Score (%)84.2
HumanLikeness - Sound-263.97Humanlike Score (%)84.2

Interactive version: theaggregate.ai/model?slug=codellama-34b-instruct-hf · How the rankings work · Data refreshed daily, snapshot 2026-07-22.