TraceBench (Low Noise, Direct): leaderboard

Metric: Accuracy (%; low observation noise; the agent submits its parameter prediction directly; controlled root-cause attribution tasks generated from three interpretable mechanical simulators (BallDrop, BounceBall, MassSlide): the agent reads time-series observations and decides whether a system parameter was altered and which one; chance 0.189, mean over three simulators). Source: arxiv.org. Saturation forecast: Rough model projection: around 2026. 4 models tracked.

Top models

#ModelScore
1gpt-5.5 (Codex, high)93.3
2claude-opus-4.6 (Claude Code, high)85.3
3gemini-3.1-pro (Gemini CLI, high)76
4minimax-m2.7 (OpenCode)29.3

Interactive version: theaggregate.ai/benchmark?slug=tracebench-low-noise-direct · How It Works · Data refreshed daily, snapshot 2026-09-26.