Terminal-Bench-LILT - Hindi: leaderboard
Metric: Pass rate (%; share of 150 trials, five per task, on the 30 Hindi-language terminal coding tasks written by native-speaker programmers, each verified by its task tests; terminus-2 harness, provider-default reasoning; failed trials caused by infrastructure were re-run). Source: arxiv.org. Saturation forecast: Around July 2027. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.8 | 52.7 |
| 2 | Gemini 3.5 Flash | 52 |
| 3 | GPT-5.5 | 51.3 |
Interactive version: theaggregate.ai/benchmark?slug=terminal-bench-lilt-hindi · How It Works · Data refreshed daily, snapshot 2026-09-29.