LiveAgentBench (No Tools): leaderboard

Metric: Share of the 374 LiveAgentBench tasks solved (%), pass@1, string-matched against closed answers, by LLMs prompted zero-shot with no tools; tasks whose attachments or media a model cannot take in count as failed; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 5 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 2.5 Pro16.85#145
2DeepSeek R19.89#245
3GPT-4o9.09#333
4Claude 3.5 Sonnet8.28#337
5Qwen 3 235B A22B7.75#304

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=liveagentbench-no-tools · How It Works · Data refreshed daily, snapshot 2026-10-11.