LongDS v1 - Education: leaderboard

Metric: Accuracy (%; macro-average over 8 education tasks of per-task turn accuracy; DSGym ReAct agent in a persistent Jupyter kernel, up to 40 steps per turn, DeepSeek-V4-Pro judge). Source: arxiv.org. Saturation forecast: Around December 2026. 5 models tracked.

Top models

#ModelScore
1GPT-5.477.92
2Claude Sonnet 4.677.29
3Kimi K2.664.98
4DeepSeek V4 Pro61.36
5Gemini 3.1 Pro (Preview)58.03

Interactive version: theaggregate.ai/benchmark?slug=longds-v1-education · How It Works · Data refreshed daily, snapshot 2026-09-25.