LoopArena - Contract Selection: leaderboard

Metric: Contract accuracy (%; 90 execution-validated questions asking which of four candidate Loop Contracts should be issued next from a frozen run state (Type I), one response per question, invalid responses counting as wrong; no Worker execution at evaluation time; provider-default reasoning). Source: arxiv.org. Saturation forecast: Around December 2026. 5 models tracked.

Top models

#ModelScore
1GPT-5.587.78
2DeepSeek V4 Flash (0731)77.78
3Claude Opus 4.876.67
4GLM-5.274.44
5Qwen 3.7 Plus72.22

Interactive version: theaggregate.ai/benchmark?slug=looparena-contract-selection · How It Works · Data refreshed daily, snapshot 2026-09-29.