Silo-Bench - Level II (Mesh): leaderboard
Metric: Success rate S (%): the share of agents whose submitted answer equals the generator's exact ground truth, averaged over the 10 Level II mesh tasks whose shards depend on their neighbours (prefix sum, moving average and the like) and all 18 configurations (team sizes 2, 5, 10, 20, 50 and 100 agents of one model, broadcast, peer-to-peer and shared-file-system protocols, each agent holding one shard of the input); higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 3 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | DeepSeek V3.1 | 35.1 | #260 |
| 2 | GPT-OSS-120B | 14.5 | #330 |
| 3 | Qwen 3 Next 80B A3B | 2.9 | #306 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=silo-bench-level-ii-mesh · How It Works · Data refreshed daily, snapshot 2026-10-11.