SWE-QA (Multi-Hop Code) - Declaration-and-Call: leaderboard

Metric: Accuracy (%) on the 4,584 Declaration-and-Call questions, which link an entity definition in one chunk to its use in another, of SWE-QA, multi-hop multiple-choice questions (four options) about the source of 12 SWE-bench Python repositories, generated with Llama-3.2-3B-Instruct; oracle setting (only the relevant code chunks), zero-shot, labels after the 66-item consensus correction; chance 25; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScore
1Llama 3.3 70B Instruct78.75
2Gemma 3 4B (IT)75.32
3Qwen 3 4B Instruct74.58
4GPT-OSS-20B67.95
5GPT-OSS-20B (Low)67.62
6GPT-OSS-20B (High)67.51
7Phi-4-mini (Reasoning)59.77
8Phi-4 Mini Instruct58.09
9Qwen 3 1.7B48.32
10DeepSeek R1 Distill Qwen 1.5B46.72
11SmolLM2-1.7B-Instruct33.92
12SmolLM2-360M-Instruct19.28

Interactive version: theaggregate.ai/benchmark?slug=swe-qa-multi-hop-code-declaration-and-call · How It Works · Data refreshed daily, snapshot 2026-10-07.