TraceBench (Low Noise, Program): leaderboard
Metric: Accuracy (%; low observation noise; the agent must submit a Python program mapping each observation to a predicted label; controlled root-cause attribution tasks generated from three interpretable mechanical simulators (BallDrop, BounceBall, MassSlide): the agent reads time-series observations and decides whether a system parameter was altered and which one; chance 0.189, mean over three simulators). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | gpt-5.5 (Codex, high) | 85.3 |
| 2 | gemini-3.1-pro (Gemini CLI, high) | 73.3 |
| 3 | claude-opus-4.6 (Claude Code, high) | 68.7 |
| 4 | minimax-m2.7 (OpenCode) | 28 |
Interactive version: theaggregate.ai/benchmark?slug=tracebench-low-noise-program · How It Works · Data refreshed daily, snapshot 2026-09-26.