MerchantBench (Hermes Agent): leaderboard

Metric: Final net assets (thousand RMB; 365-day order-level simulation of a drop-shipping store grounded in 98,843 real product records, 26 tools for sourcing, listing, pricing and cash flow; the store starts with RMB 2,000 cash and a RMB 1,000 deposit; mean of three runs; the Nous Research Hermes Agent default configuration with its code, planning, memory and skill tools). Source: arxiv.org. Saturation forecast: Around October 2027. 8 models tracked.

Top models

#ModelScore
1Qwen 3.7 Max59.46
2GPT-5.6 Sol52.93
3GLM-5.242.32
4Claude Opus 4.835.56
5Qwen 3.7 Plus29.42
6DeepSeek V4 Flash24.69
7Kimi K2.623.96
8DeepSeek V4 Pro16.71

Interactive version: theaggregate.ai/benchmark?slug=merchantbench-hermes-agent · How It Works · Data refreshed daily, snapshot 2026-09-29.