ResearchGym - Continual Learning: leaderboard

Metric: Best-of-3 score normalised to the task's reference (SOTA) result, SOTA = 100. Source: anikethh.github.io. Saturation forecast: Rough model projection: around 2026. 3 models tracked.

Top models

#ModelScore
1claude-opus-4.5 (Claude Code)98.5
2gpt-5.2-codex (Codex, xhigh)97.9
3RG-Agent + gpt-5 (high)94.3

Interactive version: theaggregate.ai/benchmark?slug=researchgym-continual-learning · How It Works · Data refreshed daily, snapshot 2026-09-26.