Terminal Bench — leaderboard

Terminal and command-line interaction tasks for evaluating agent performance.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 9 models tracked.

Top models

#ModelScore
1GPT-5.264.9
2Gemini 3 Pro64.3
3Claude Opus 4.563.1
4Kimi K2 (Thinking)35.7
5Gemini 2.5 Pro (Preview 06-05)32.6
6Grok 427.2
7Grok Code Fast 125.8
8GPT-OSS-120B18.7

Interactive version: theaggregate.ai/benchmark?slug=terminal-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.