TERMS-Bench - Surplus Efficiency: leaderboard

Metric: Surplus efficiency on feasible episodes (SE+, maximum 1; loss-making agreements are not clipped and count below 0): the agent's utility divided by the zone-of-possible-agreement width, averaged over 1,800 seeded episodes per agent (six counterpart families, three regimes), deterministic decoding; higher is better. Source: arxiv.org. Saturation forecast: Around September 2027. 13 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (Max)0.69
2GLM-5.10.69
3Claude Opus 4.7 (Max)0.66
4Gemma 4 31B (IT) (Thinking)0.64
5Gemini 3.1 Pro (Preview) (High)0.64
6DeepSeek V4 Pro (Max)0.62
7GPT-5.5 (xHigh)0.61
8Qwen 3.6 Plus0.6
9Grok 4.20 (Reasoning)0.6
10Kimi K2.60.6
11GPT-5.4 (xHigh)0.53
12Seed 2.0 Pro (High)0.52
13GPT-4o Mini0.19

Interactive version: theaggregate.ai/benchmark?slug=terms-bench-surplus-efficiency · How It Works · Data refreshed daily, snapshot 2026-10-07.