Long Code Arena — leaderboard

JetBrains long-context code benchmark: 5 tasks (library code generation, CI repair, commit message generation, bug localization, module summarization) testing real-world repository-scale understanding.

Metric: Mean Score (%). Source: huggingface.co. Status: saturated. 9 models tracked.

Top models

#ModelScore
1O195.7
2Claude 3.5 Sonnet83.8
3DeepSeek R179.6
4GPT-4o70.2
5Gemini 1.5 Pro58.1
6Llama 3.1 405B46.8
7Claude 3 Haiku42.3
8Llama 3.1 70B28.8
9Llama 3.1 8B11.7

Interactive version: theaggregate.ai/benchmark?slug=long-code-arena · How the rankings work · Data refreshed daily, snapshot 2026-07-22.