ArkEval - File Localization: leaderboard
Metric: Macro-F1 (%) per issue between the model's predicted repair-relevant file set and the curated gold set from the developer patch, over the 502 issues, after a shared ArkTS-aware embedding retrieval of ten candidate files and the model's two-pass scope refinement; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 8 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT-5.6 Sol | 34.87 | #16 |
| 2 | Qwen 3.7 Max | 33.24 | #71 |
| 3 | Kimi K2.7 Code | 32.77 | #81 |
| 4 | GLM-5.2 | 31.75 | #69 |
| 5 | DeepSeek V4 Pro | 30.53 | #96 |
| 6 | MiMo-V2.5-Pro | 29.14 | #152 |
| 7 | MiniMax-M3 | 25.14 | #129 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=arkeval-file-localization · How It Works · Data refreshed daily, snapshot 2026-10-11.