CoT-Control (OpenAI) — leaderboard
Measures whether reasoning models can suppress or hide their chain-of-thought - a key AI safety property. 13,000+ tasks from GPQA, MMLU-Pro, HLE, BFCL, and SWE-Bench Verified. Frontier models score 0.1%–15.4%, meaning they largely cannot control their CoT, which is good for monitorability.
Source: openai.com.
Interactive version: theaggregate.ai/benchmark?slug=cot-control-openai · How the rankings work · Data refreshed daily, snapshot 2026-07-22.