SMH-Bench (Interactive Agent): leaderboard

Metric: Success rate (%), instance-weighted over 1,100 tasks in 7 capability categories; Environment-Interactive Agent: a ReAct loop that starts from partial observability and queries and controls the HomeEnv simulator through tools under a call budget; final home state checked against the goal with untouched devices preserved; clarification and query answers judged by GPT-5; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)85.2
2Claude Sonnet 4.684.1
3GPT-5.479.6
4Qwen 3.5 397B A17B77.8
5Claude Haiku 4.576.2
6DeepSeek V3.2 (Thinking)75.5
7GPT-5.4 Mini75.3
8GLM-570.6
9Qwen 3.5 Plus70
10Qwen 3.5 9B64.3
11Qwen 3.5 4B56.5

Interactive version: theaggregate.ai/benchmark?slug=smh-bench-interactive-agent · How It Works · Data refreshed daily, snapshot 2026-09-29.