SMH-Bench (Direct Reasoning) - State Query: leaderboard

Metric: Success rate (%) on Environment-Grounded Query (TC7, 100 non-mutating inspection and aggregation questions, GPT-5 judged); Direct Reasoning: the full serialized home state (rooms, devices, current values, service schemas) and the request in one prompt, one structured JSON answer replayed in the HomeEnv simulator; final home state checked against the goal with untouched devices preserved; clarification and query answers judged by GPT-5; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 13 models tracked.

Top models

#ModelScore
1GPT-5.495
2Gemini 3.1 Pro (Preview)94
3GPT-5.4 Mini89
4DeepSeek V3.2 (Thinking)89
5Claude Sonnet 4.685
6Qwen 3.5 397B A17B84
7Claude Haiku 4.579
8GLM-578
9Qwen 3.5 9B77
10Qwen 3.5 4B75
11Qwen 3.5 Plus69
12MiniMax-M2.763

Interactive version: theaggregate.ai/benchmark?slug=smh-bench-direct-reasoning-state-query · How It Works · Data refreshed daily, snapshot 2026-09-29.