Ko-AgentBench - L3 Sequential Tool Reasoning — leaderboard

Metric: PSM (%). Source: huggingface.co. 15 models tracked.

Top models

#ModelScore
1GPT-4o100
2Claude Sonnet 4.5100
3Claude Haiku 4.5100
4nova-2-lite-v1100
5Grok 4.1 Fast100
6Qwen 3 Next 80B A3B100
7GLM-4.6V96.67
8GPT-4o Mini91.67
9Nova Lite85
10DeepSeek V3.180
11GPT-556.67
12GPT-5 Mini55
13Gemini 2.5 Flash55
14Gemini 2.5 Pro26.67
15Gemini 2.5 Flash Lite20

Interactive version: theaggregate.ai/benchmark?slug=ko-agentbench-l3-sequential-tool-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.