ArkEval: leaderboard
Metric: Pass@1 (%): share of instances whose patch compiles and passes the executable reproduction test (and the existing regression suite where one exists), on all 502 ArkTS/OpenHarmony issue instances from nine repositories, one patch per instance from the ArkFix pipeline (each model's own file localization, then patch-only repair at temperature 0 with at most 50 agent steps, no compiler, test or device feedback; official-documentation RAG off); higher is better. Source: arxiv.org. Saturation forecast: Around 2028. 8 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | GPT-5.6 Sol | 19.32 | #16 |
| 2 | GLM-5.2 | 8.37 | #69 |
| 3 | MiniMax-M3 | 7.17 | #129 |
| 4 | Qwen 3.7 Max | 6.18 | #71 |
| 5 | MiMo-V2.5-Pro | 5.98 | #152 |
| 6 | DeepSeek V4 Pro | 5.18 | #96 |
| 7 | Kimi K2.7 Code | 3.19 | #81 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=arkeval · How It Works · Data refreshed daily, snapshot 2026-10-11.