ArkEval - Compile@1: leaderboard
Metric: Compile@1 (%): share of instances whose patch completes the configured hvigor ArkTS build, unconditioned on earlier gates, on all 502 ArkTS/OpenHarmony issue instances from nine repositories, one patch per instance from the ArkFix pipeline (each model's own file localization, then patch-only repair at temperature 0 with at most 50 agent steps, no compiler, test or device feedback; official-documentation RAG off); higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 8 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GLM-5.2 | 54.38 | #69 |
| 2 | MiMo-V2.5-Pro | 53.39 | #152 |
| 3 | MiniMax-M3 | 44.42 | #129 |
| 4 | DeepSeek V4 Pro | 41.43 | #96 |
| 5 | Qwen 3.7 Max | 37.25 | #71 |
| 6 | GPT-5.6 Sol | 34.26 | #16 |
| 7 | Kimi K2.7 Code | 12.35 | #81 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=arkeval-compile-1 · How It Works · Data refreshed daily, snapshot 2026-10-11.