Terminal-Bench 4.0: leaderboard

Sixty-six containerized terminal tasks weighted toward science-adjacent and frontier engineering work, each agent run five times. Relative to 2.1 it lengthens timeouts and raises RAM and CPU on selected tasks, so a failure is more likely the agent's than the harness's.

Metric: Accuracy (%). Source: www.tbench.ai. Status: years away from saturation. 14 models tracked.

Top models

#ModelScore
1GPT-658.2
2Claude Fable 5.157.9
3Claude Opus 551.8
4Claude Fable 544.5
5GLM-5.341.8
6GPT-5.6 Sol37.3
7Claude Opus 4.823.6
8GPT-5.6 Terra21.5
9Grok 4.620.3
10Gemini 3.8 Flash19.1
11GPT-5.6 Luna17.3
12Grok 4.512.4
13Claude Sonnet 512.4
14Gemini 3.7 Flash11.2

Interactive version: theaggregate.ai/benchmark?slug=terminal-bench-4-0 · How It Works · Data refreshed daily, snapshot 2026-09-05.