ResearchGym - Improving Replay Buffers: leaderboard
Metric: Best-of-3 score normalised to the task's reference (SOTA) result, SOTA = 100. Source: anikethh.github.io. Saturation forecast: Rough model projection: around 2026. 3 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | gpt-5.2-codex (Codex, xhigh) | 49.7 |
| 2 | RG-Agent + gpt-5 (high) | 34.3 |
| 3 | claude-opus-4.5 (Claude Code) | 21.7 |
Interactive version: theaggregate.ai/benchmark?slug=researchgym-improving-replay-buffers · How It Works · Data refreshed daily, snapshot 2026-09-26.