ClarifyCodeBench: leaderboard

Metric: Pass@1 (%) on the hidden LiveCodeBench v6 tests of the code written after interactive clarification of an ambiguous requirement (419 tasks with one to three deleted requirement details; a GPT-4o judge matches each question to the annotated key questions and the environment returns the deleted detail; at most 6 rounds, temperature 0). Source: arxiv.org. Saturation forecast: Around January 2028. 7 models tracked.

Top models

#ModelScore
1GPT-541.1
2DeepSeek V3.2 (Non-reasoning)39.1
3Claude Sonnet 4.538.2
4Claude Sonnet 4.5 (Thinking)34.3
5Gemini 2.5 Flash (Non-reasoning)29.8
6Qwen 3 235B A22B (Non-reasoning)27.7
7GPT-4o27.2

Interactive version: theaggregate.ai/benchmark?slug=clarifycodebench · How It Works · Data refreshed daily, snapshot 2026-09-29.