AMA-Bench — leaderboard

Agent and memory benchmark spanning text-to-SQL, software, web, games, embodied AI, and open-world question answering tasks.

Metric: Average Score (%). Source: huggingface.co. Status: years away from saturation. 24 models tracked.

Top models

#ModelScore
1GPT-5.269.83
2GPT-5 Mini65.64
3Qwen 3 32B50.17
4Gemini 2.5 Flash49.22
5Qwen 3 14B45.24
6Qwen2.5-14B-Instruct-1M45.07
7Claude 3.5 Haiku43.12
8Qwen 3 8B39.89

Interactive version: theaggregate.ai/benchmark?slug=ama-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.