AsgardBench (Text-Only): leaderboard

Metric: Task success rate (%) with text-only observations (no images, simple feedback), on AsgardBench's 108 AI2-THOR household tasks (12 task types with varied object states and placements) in which the agent issues high-level actions with navigation and grasping abstracted away and must adapt its plan to what it observes, receiving only simple success or failure feedback; temperature 0, median over independent runs; higher is better. Source: arxiv.org. Saturation forecast: Around July 2027. 9 models tracked.

Top models

#ModelScoreOverall rank
1Claude Opus 4.5 (High)36.1#79 (Claude Opus 4.5)
2Gemini 3 Pro (Preview) (High)35.2#64 (Gemini 3 Pro (Preview))
3GPT-5.2 (High)25#105 (GPT-5.2)
4Kimi K2.521.9#139
5Qwen 3 VL 235B A22B (Thinking)13.9#228 (Qwen 3 VL 235B A22B)
6GLM-4.6V8.3#309
7Mistral Large 37.9#388
8Llama 4 Maverick Instruct FP86.7#403
9GPT-4o4.2#333

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=asgardbench-text-only · How It Works · Data refreshed daily, snapshot 2026-10-11.