ArkEval: leaderboard

Metric: Pass@1 (%): share of instances whose patch compiles and passes the executable reproduction test (and the existing regression suite where one exists), on all 502 ArkTS/OpenHarmony issue instances from nine repositories, one patch per instance from the ArkFix pipeline (each model's own file localization, then patch-only repair at temperature 0 with at most 50 agent steps, no compiler, test or device feedback; official-documentation RAG off); higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 8 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.6 Sol19.32#16
2GLM-5.28.37#69
3MiniMax-M37.17#129
4Qwen 3.7 Max6.18#71
5MiMo-V2.5-Pro5.98#152
6DeepSeek V4 Pro5.18#96
7Kimi K2.7 Code3.19#81

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=arkeval · How It Works · Data refreshed daily, snapshot 2026-10-11.