Terminal-Bench-LILT - Chinese: leaderboard

Metric: Pass rate (%; share of 150 trials, five per task, on the 30 Chinese-language terminal coding tasks written by native-speaker programmers, each verified by its task tests; terminus-2 harness, provider-default reasoning; failed trials caused by infrastructure were re-run). Source: arxiv.org. Saturation forecast: Around January 2027. 3 models tracked.

Top models

#ModelScore
1Claude Opus 4.864.7
2Gemini 3.5 Flash64.7
3GPT-5.553.3

Interactive version: theaggregate.ai/benchmark?slug=terminal-bench-lilt-chinese · How It Works · Data refreshed daily, snapshot 2026-09-29.