Market-Bench: leaderboard
Metric: Cumulative profit (simulation currency units, signed; each agent starts with 22,500 in funds); all 20 models compete as retailer agents in one simulated supply-chain market (sealed-bid procurement auctions, retail pricing and slogans; 6 steps, temperature 0), mean over 10 runs; higher is better. Source: arxiv.org. Saturation forecast: Around March 2027. 20 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 2.5 Pro | 36589 |
| 2 | Gemini 2.5 Flash | 26104 |
| 3 | O3 | 13368 |
| 4 | Claude Sonnet 4.5 | 10619 |
| 5 | GPT-4o | 7619 |
| 6 | Phi-4 | 7565 |
| 7 | Qwen 2.5 VL 72B Instruct | 3402 |
| 8 | Qwen 2.5 VL 32B Instruct | 1409 |
| 9 | ERNIE 4.5 300B A47B | 1360 |
| 10 | InternLM3-8B-Instruct | 1022 |
| 11 | DeepSeek V3.2 | 616 |
Interactive version: theaggregate.ai/benchmark?slug=market-bench · How It Works · Data refreshed daily, snapshot 2026-10-07.