Silo-Bench: leaderboard

Metric: Success rate S (%): the share of agents whose submitted answer equals the generator's exact ground truth, averaged over all 30 tasks and all 18 configurations (team sizes 2, 5, 10, 20, 50 and 100 agents of one model, broadcast, peer-to-peer and shared-file-system protocols, each agent holding one shard of the input); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 3 models tracked.

Top models

#ModelScoreOverall rank
1DeepSeek V3.136.9#260
2GPT-OSS-120B16.9#330
3Qwen 3 Next 80B A3B8.2#306

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=silo-bench · How It Works · Data refreshed daily, snapshot 2026-10-11.