FORTIS - Skill-Grounded Tool Selection Fail Rate: leaderboard

Metric: Fail rate (%), the share of over-privileged or no-action outcomes, on FORTIS Task 2 (skill-grounded tool selection, 1,543 queries: the model gets the assigned skill document and the full tool inventory and must pick the minimum-privilege feasible tool set), temperature 0; lower is better. Source: arxiv.org. Saturation forecast: Not forecast. 10 models tracked.

Top models

#ModelScore
1Qwen 3.6 Max45.2
2Claude Opus 4.747.4
3Gemini 3.1 Pro (Preview)49.2
4DeepSeek V4 Flash53
5Kimi K2.653.7
6Claude Sonnet 4.656.8
7GPT-5.4 Mini59.3
8Gemini 3 Flash61.2
9GPT-5.562.5
10GPT-5.466.6

Interactive version: theaggregate.ai/benchmark?slug=fortis-skill-grounded-tool-selection-fail-rate · How It Works · Data refreshed daily, snapshot 2026-10-07.