ResearchGym - Continual Learning: leaderboard
Metric: Best-of-3 score normalised to the task's reference (SOTA) result, SOTA = 100. Source: anikethh.github.io. Saturation forecast: Rough model projection: around 2026. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | claude-opus-4.5 (Claude Code) | 98.5 |
| 2 | gpt-5.2-codex (Codex, xhigh) | 97.9 |
| 3 | RG-Agent + gpt-5 (high) | 94.3 |
Interactive version: theaggregate.ai/benchmark?slug=researchgym-continual-learning · How It Works · Data refreshed daily, snapshot 2026-09-26.