SEAL - Agentic Tool Use (Enterprise) — leaderboard

Scale SEAL evaluation of multi-tool reasoning in enterprise scenarios with 485 prompts testing complex tool composition.

Metric: Score. Source: scale.com. Status: saturation imminent. 35 models tracked.

Top models

#ModelScore
1O1 (2024-12-17)70.14
2Gemini 2.5 Pro (03-25)68.75
3O1 Pro67.01
4O1 Preview66.43
5DeepSeek R165.27
6O3 Mini (High)65.27
7Claude 3.7 Sonnet (Thinking)65.27
8Claude 3.7 Sonnet (20250219)64.93
9O3 Mini (Medium)64.93
10GPT-4o (2024-05-13)64.58
11DeepSeek V3 (0324)64.23
12GPT-4.5 (Preview)63.76
13Gemini 2.0 Flash (Thinking)63.19
14DeepSeek V362.5
15Gemini 2.0 Pro (Preview 02-05)61.45

Interactive version: theaggregate.ai/benchmark?slug=seal-agentic-tool-use-enterprise · How the rankings work · Data refreshed daily, snapshot 2026-07-22.