ArkEval - Compile@1: leaderboard

Metric: Compile@1 (%): share of instances whose patch completes the configured hvigor ArkTS build, unconditioned on earlier gates, on all 502 ArkTS/OpenHarmony issue instances from nine repositories, one patch per instance from the ArkFix pipeline (each model's own file localization, then patch-only repair at temperature 0 with at most 50 agent steps, no compiler, test or device feedback; official-documentation RAG off); higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 8 models tracked.

Top models

#ModelScoreOverall rank
1GLM-5.254.38#69
2MiMo-V2.5-Pro53.39#152
3MiniMax-M344.42#129
4DeepSeek V4 Pro41.43#96
5Qwen 3.7 Max37.25#71
6GPT-5.6 Sol34.26#16
7Kimi K2.7 Code12.35#81

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=arkeval-compile-1 · How It Works · Data refreshed daily, snapshot 2026-10-11.