AsgardBench: leaderboard

Metric: Task success rate (%) with image observations (the previous and current views), on AsgardBench's 108 AI2-THOR household tasks (12 task types with varied object states and placements) in which the agent issues high-level actions with navigation and grasping abstracted away and must adapt its plan to what it observes, receiving only simple success or failure feedback; temperature 0, median over independent runs; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 9 models tracked.

Top models

#ModelScoreOverall rank
1Claude Opus 4.5 (High)76.5#79 (Claude Opus 4.5)
2Gemini 3 Pro (Preview) (High)72.5#64 (Gemini 3 Pro (Preview))
3GPT-5.2 (High)71.3#105 (GPT-5.2)
4Kimi K2.568.8#139
5Qwen 3 VL 235B A22B (Thinking)32.4#228 (Qwen 3 VL 235B A22B)
6GLM-4.6V30.6#309
7GPT-4o23.4#333
8Mistral Large 37.4#388
9Llama 4 Maverick Instruct FP85.8#403

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=asgardbench · How It Works · Data refreshed daily, snapshot 2026-10-11.