FORTIS - Skill-Grounded Tool Selection: leaderboard

Metric: Exact-match rate (%) with the minimum-privilege tool set on FORTIS Task 2 (skill-grounded tool selection, 1,543 queries: the model gets the assigned skill document and the full tool inventory and must pick the minimum-privilege feasible tool set), temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 10 models tracked.

Top models

#ModelScore
1Claude Opus 4.726.1
2Qwen 3.6 Max25.9
3Kimi K2.623.3
4Gemini 3.1 Pro (Preview)22.7
5Claude Sonnet 4.622.4
6DeepSeek V4 Flash20.7
7GPT-5.520.4
8GPT-5.4 Mini19.7
9Gemini 3 Flash16.9
10GPT-5.415.2

Interactive version: theaggregate.ai/benchmark?slug=fortis-skill-grounded-tool-selection · How It Works · Data refreshed daily, snapshot 2026-10-07.