SWE-QA (Multi-Hop Code) - Interacting Entities: leaderboard

Metric: Accuracy (%) on the 4,488 Interacting-Entity questions, which follow two entities interacting across three chunks, of SWE-QA, multi-hop multiple-choice questions (four options) about the source of 12 SWE-bench Python repositories, generated with Llama-3.2-3B-Instruct; oracle setting (only the relevant code chunks), zero-shot, labels after the 66-item consensus correction; chance 25; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 15 models tracked.

Top models

#ModelScore
1Llama 3.3 70B Instruct69.98
2Gemma 3 4B (IT)68.6
3Qwen 3 4B Instruct65.41
4GPT-OSS-20B53.83
5GPT-OSS-20B (Low)53.43
6GPT-OSS-20B (High)53.16
7Phi-4-mini (Reasoning)50.35
8Phi-4 Mini Instruct47.9
9Qwen 3 1.7B43.18
10DeepSeek R1 Distill Qwen 1.5B39.14
11SmolLM2-1.7B-Instruct29.94
12SmolLM2-360M-Instruct20.61

Interactive version: theaggregate.ai/benchmark?slug=swe-qa-multi-hop-code-interacting-entities · How It Works · Data refreshed daily, snapshot 2026-10-07.