TERMS-Bench - Surplus Efficiency: leaderboard
Metric: Surplus efficiency on feasible episodes (SE+, maximum 1; loss-making agreements are not clipped and count below 0): the agent's utility divided by the zone-of-possible-agreement width, averaged over 1,800 seeded episodes per agent (six counterpart families, three regimes), deterministic decoding; higher is better. Source: arxiv.org. Saturation forecast: Around September 2027. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.6 (Max) | 0.69 |
| 2 | GLM-5.1 | 0.69 |
| 3 | Claude Opus 4.7 (Max) | 0.66 |
| 4 | Gemma 4 31B (IT) (Thinking) | 0.64 |
| 5 | Gemini 3.1 Pro (Preview) (High) | 0.64 |
| 6 | DeepSeek V4 Pro (Max) | 0.62 |
| 7 | GPT-5.5 (xHigh) | 0.61 |
| 8 | Qwen 3.6 Plus | 0.6 |
| 9 | Grok 4.20 (Reasoning) | 0.6 |
| 10 | Kimi K2.6 | 0.6 |
| 11 | GPT-5.4 (xHigh) | 0.53 |
| 12 | Seed 2.0 Pro (High) | 0.52 |
| 13 | GPT-4o Mini | 0.19 |
Interactive version: theaggregate.ai/benchmark?slug=terms-bench-surplus-efficiency · How It Works · Data refreshed daily, snapshot 2026-10-07.