Cangjie-bench - Code Generation (ILA-agent): leaderboard

Metric: Accuracy (%; share of solutions passing all private unit tests; 155 HumanEval problems written in the Cangjie language; ILA-agent: the model explores the official Cangjie documentation and verifies code with the compiler through tools, at most 15 turns). Source: arxiv.org. Saturation forecast: Around December 2026. 3 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.581.94
2Qwen 3 Max73.55
3DeepSeek V3.263.23

Interactive version: theaggregate.ai/benchmark?slug=cangjie-bench-code-generation-ila-agent · How It Works · Data refreshed daily, snapshot 2026-09-26.