RepoQA: leaderboard

RepoQA: Evaluates software-engineering agents on realistic issue resolution, repository navigation, testing, or maintenance workflows.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturated. 33 models tracked.

Top models

#ModelScore
1Gemini 1.5 Pro90.6
2GPT-4o (2024-05-13)90.6
3Claude 3 Opus (20240229)90.6
4Gemini 1.5 Flash90
5Claude 3 Sonnet (20240229)87.4
6DeepSeek V2 Chat83.4
7Llama 3 70B Instruct82.2
8Claude 3 Haiku (20240307)81.8
9c4ai-command-r-plus78.4
10GPT-4 Turbo76.4
11Mixtral 8x7B Instruct (v0.1)68
12Mixtral 8x22B Instruct (v0.1)67.8
13Qwen 1.5 72B Chat67
14Phi-3-medium-128k-instruct63.2
15Mistral 7B Instruct (v0.3)62

Interactive version: theaggregate.ai/benchmark?slug=repoqa · How It Works · Data refreshed daily, snapshot 2026-09-05.