ArkEval - File Localization: leaderboard

Metric: Macro-F1 (%) per issue between the model's predicted repair-relevant file set and the curated gold set from the developer patch, over the 502 issues, after a shared ArkTS-aware embedding retrieval of ten candidate files and the model's two-pass scope refinement; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 8 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.6 Sol34.87#16
2Qwen 3.7 Max33.24#71
3Kimi K2.7 Code32.77#81
4GLM-5.231.75#69
5DeepSeek V4 Pro30.53#96
6MiMo-V2.5-Pro29.14#152
7MiniMax-M325.14#129

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=arkeval-file-localization · How It Works · Data refreshed daily, snapshot 2026-10-11.