YC-Bench — leaderboard

Long-horizon agentic startup simulation: models run an AI company for 12 months managing tasks, employees, clients, and cash flow. Score = final net worth averaged over 3 seeds.

Metric: Net Worth ($K). Source: collinear-ai.github.io. Status: saturation imminent. 29 models tracked.

Top models

#ModelScore
1Claude Fable 51977.6
2Claude Opus 4.71714.5
3Claude Opus 4.81711.4
4GLM-5.11510.8
5Claude Opus 4.61269.7
6MiMo-V2.51220.9
7GLM-51208.2
8GPT-5.51206
9Claude Sonnet 51163.5
10DeepSeek V4 Pro1066.4
11GLM-5.21013.2
12GPT-5.41000.8
13MiniMax-M3999.5
14Gemini 3.5 Flash987
15Qwen 3.6 Plus788

Interactive version: theaggregate.ai/benchmark?slug=yc-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.