TERMS-Bench — leaderboard

Negotiation benchmark for LLM agents bargaining over terms under changing utility, urgency, and no-deal regimes, reporting mean utility and agreement metrics.

Metric: Mean Utility. Source: terms-bench.github.io. Status: saturation imminent. 15 models tracked.

Top models

#ModelScore
1GLM-5.111.7
2Claude Opus 4.611.62
3Claude Opus 4.711.09
4Gemma 4 31B10.64
5Gemini 3.1 Pro (Preview)10.58
6DeepSeek V4 Pro10.53
7Qwen 3.6 Plus10.3
8Kimi K2.610.17
9GPT-5.510.15
10Grok 4.2010.03
11Seed 2.0 Pro8.61
12GPT-4o Mini3.37

Interactive version: theaggregate.ai/benchmark?slug=terms-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.