AppWorld Normal — leaderboard

Interactive coding agent benchmark simulating 9 everyday apps (email, calendar, Spotify, etc.) with 457 APIs and 750+ test scenarios. Normal difficulty split.

Metric: Task Goal Completion (%). Source: appworld.dev. 7 models tracked.

Top models

#ModelScore
1Qwen 3 14B86.9
2GPT-4.173.2
3Qwen 2.5 32B72.6
4GPT-4o68.5
5GPT-4 Turbo32.7
6Llama 3 70B24.4

Interactive version: theaggregate.ai/benchmark?slug=appworld-normal · How the rankings work · Data refreshed daily, snapshot 2026-07-22.