MiniAppBench - Lifestyle: leaderboard

Metric: Pass rate (%) on the 32 Lifestyle tasks, MiniAppBench's 500 interactive-HTML MiniApp tasks distilled from real user queries; the model writes a self-contained HTML/JS app in one generation under its official decoding settings, and the MiniAppEval browser agent (Gemini-3-Pro-Preview driving Playwright) scores Intention, Static and Dynamic fidelity from 0 to 1; a task passes when the minimum of the three scores is at least 0.8; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 16 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.282.35#105
2GPT-5.164.71#131
3Claude Opus 4.556.52#79
4Gemini 3 Pro (Preview)55.56#64
5GLM-4.748.39#185
6Claude Sonnet 4.544.83#138
7Gemini 3 Flash41.38#93
8MiMo-V2-Flash36.36#359
9Grok 4.1 Fast (Reasoning)25.93#208 (Grok 4.1 Fast)
10MiniMax-M2.119.23#283
11Kimi K218.52#236
12Qwen 3 Coder 480B A35B Instruct11.11#302
13Qwen 3 235B A22B10.34#304
14GLM 4.5 Air10.34#377
15Qwen 3 32B3.7#424

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=miniappbench-lifestyle · How It Works · Data refreshed daily, snapshot 2026-10-11.