Terminal Bench — leaderboard
Terminal and command-line interaction tasks for evaluating agent performance.
Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 | 64.9 |
| 2 | Gemini 3 Pro | 64.3 |
| 3 | Claude Opus 4.5 | 63.1 |
| 4 | Kimi K2 (Thinking) | 35.7 |
| 5 | Gemini 2.5 Pro (Preview 06-05) | 32.6 |
| 6 | Grok 4 | 27.2 |
| 7 | Grok Code Fast 1 | 25.8 |
| 8 | GPT-OSS-120B | 18.7 |
Interactive version: theaggregate.ai/benchmark?slug=terminal-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.