TraceBench (Low Noise, Direct): leaderboard
Metric: Accuracy (%; low observation noise; the agent submits its parameter prediction directly; controlled root-cause attribution tasks generated from three interpretable mechanical simulators (BallDrop, BounceBall, MassSlide): the agent reads time-series observations and decides whether a system parameter was altered and which one; chance 0.189, mean over three simulators). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | gpt-5.5 (Codex, high) | 93.3 |
| 2 | claude-opus-4.6 (Claude Code, high) | 85.3 |
| 3 | gemini-3.1-pro (Gemini CLI, high) | 76 |
| 4 | minimax-m2.7 (OpenCode) | 29.3 |
Interactive version: theaggregate.ai/benchmark?slug=tracebench-low-noise-direct · How It Works · Data refreshed daily, snapshot 2026-09-26.