SkillsBench — leaderboard

SkillsBench (BenchFlow) evaluates how effectively agentic models use modular Skills — folders of instructions, scripts, and resources — to complete tasks across diverse domains; scored by task completion with the Skills available (the benchmark also reports skill-lift, the with-minus-without-Skills gain).

Metric: With-skills score (%). Source: www.skillsbench.ai. Status: saturation imminent. 21 models tracked.

Top models

#ModelScore
1GPT-5.567.3
2Claude Opus 4.761.2
3Gemini 3.1 Pro (Preview)60.8
4GLM-5.158.4
5Gemini 3 Flash54.6
6Claude Opus 4.854.1
7Kimi K2.654
8MiniMax-M353
9GPT-5.251.7
10Claude Opus 4.650.2
11DeepSeek V4 Pro50.1
12Claude Opus 4.549
13Gemini 3.5 Flash48.2
14Claude Sonnet 4.647.2
15DeepSeek V4 Flash44.7

Interactive version: theaggregate.ai/benchmark?slug=skillsbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.