CanAiCode — leaderboard

Benchmarks smaller code-focused LLMs on text-to-code performance across multiple languages.

Metric: Junior-v2 Python Pass Rate (%). Source: huggingface.co. Status: saturated. 189 models tracked.

Top models

#ModelScore
1Llama 3.1 8B Instruct100
2Claude 3.5 Sonnet (20240620)100
3GPT-4 Turbo100
4Claude 3 Opus (20240229)100
5Claude 3 Sonnet (20240229)100
6GPT-4 (0613)100
7GPT-4 Preview (1106)100
8Phi-3-small-8k-instruct100
9Phi-3-medium-128k-instruct100
10GPT-3.5 Turbo (0301)100
11GPT-4 Preview (0125)100
12DeepSeek Coder V2 Lite Instruct100
13Nxcode-CQ-7B-orpo100
14Gemma 2 27B (IT)98.9
15GPT-4o (2024-05-13)98.9

Interactive version: theaggregate.ai/benchmark?slug=canaicode · How the rankings work · Data refreshed daily, snapshot 2026-07-22.