TS-Skill (Free-Form): leaderboard

Metric: Native-form score (0-1; free-form answers parsed by GPT-5.4 and scored with answer-type comparators plus a GPT-5.4 rationale judge, mean over 3,000 questions). Source: arxiv.org. Saturation forecast: Around October 2026. 9 models tracked.

Top models

#ModelScore
1GPT-4o0.67
2O3 Mini0.61
3Claude Sonnet 4.50.6
4Qwen 3 32B (Thinking)0.55
5Qwen 2.5 14B Instruct0.51
6Qwen 2.5 VL 7B0.5

Interactive version: theaggregate.ai/benchmark?slug=ts-skill-free-form · How It Works · Data refreshed daily, snapshot 2026-09-25.