FinSkillBench (Curated Skills): leaderboard

Metric: Task score (0-1, native; with human-authored skill and tool-script packages mounted; over FinSkillBench's 12 investment-management subtasks (portfolio construction, risk management, fundamental analysis), unweighted mean, hidden ground truth and task-specific verifiers). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1GPT-4.10.74
2Gemini 2.5 Pro0.68
3Claude Sonnet 4.60.66
4DeepSeek V3.20.66
5Gemini 3.1 Flash Lite0.61
6GPT-5.40.52
7GLM-5.10.28
8Grok 40.08
9Phi-40

Interactive version: theaggregate.ai/benchmark?slug=finskillbench-curated-skills · How It Works · Data refreshed daily, snapshot 2026-09-26.