LoopArena - Contract Selection: leaderboard
Metric: Contract accuracy (%; 90 execution-validated questions asking which of four candidate Loop Contracts should be issued next from a frozen run state (Type I), one response per question, invalid responses counting as wrong; no Worker execution at evaluation time; provider-default reasoning). Source: arxiv.org. Saturation forecast: Around December 2026. 5 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.5 | 87.78 |
| 2 | DeepSeek V4 Flash (0731) | 77.78 |
| 3 | Claude Opus 4.8 | 76.67 |
| 4 | GLM-5.2 | 74.44 |
| 5 | Qwen 3.7 Plus | 72.22 |
Interactive version: theaggregate.ai/benchmark?slug=looparena-contract-selection · How It Works · Data refreshed daily, snapshot 2026-09-29.