RepoQA — leaderboard

RepoQA: Evaluates software-engineering agents on realistic issue resolution, repository navigation, testing, or maintenance workflows.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturated. 33 models tracked.

Top models

#ModelScore
1Gemini 1.5 Pro90.6
2Claude 3 Opus90.6
3GPT-4o (2024-05-13)90.6
4Gemini 1.5 Flash90
5Claude 3 Sonnet (20240229)87.4
6DeepSeek V2 Chat83.4
7Llama 3 70B Instruct82.2
8Claude 3 Haiku81.8
9c4ai-command-r-plus78.4
10GPT-4 Turbo76.4
11Mixtral 8x22B Instruct (v0.1)67.8
12Qwen 1.5 72B Chat67
13Mistral 7B Instruct (v0.3)62
14GPT-3.5 Turbo60.4
15Llama 3 8B Instruct53.6

Interactive version: theaggregate.ai/benchmark?slug=repoqa · How the rankings work · Data refreshed daily, snapshot 2026-07-22.