SEAL - Agentic Tool Use (Chat) — leaderboard

Scale SEAL evaluation of tool use in conversational settings, testing ability to select and chain tools for user requests.

Metric: Score. Source: scale.com. Status: saturation imminent. 35 models tracked.

Top models

#ModelScore
1O3 Mini (High)63.45
2Gemini 2.5 Pro (03-25)62.43
3O3 Mini (Medium)62.42
4O1 Pro61.42
5DeepSeek R160.91
6O1 (2024-12-17)60.41
7DeepSeek V358.55
8Gemini 2.0 Pro (Preview 02-05)57.86
9Gemini 2.0 Flash (Thinking)57.36
10GPT-4o (2024-08-06)56.85
11GPT-4.5 (Preview)56.34
12Claude 3.7 Sonnet (20250219)56.25
13Claude 3.5 Sonnet (20240620)56.06
14Claude 3.7 Sonnet (Thinking)55.32
15O1 Preview55.1

Interactive version: theaggregate.ai/benchmark?slug=seal-agentic-tool-use-chat · How the rankings work · Data refreshed daily, snapshot 2026-07-22.