CAR-bench — leaderboard

Context-aware reasoning benchmark for agents, focused on following task context, constraints, and environment feedback during multi-step work.

Metric: Avg Pass^3 (%). Source: car-bench.github.io. Status: saturation imminent. 10 models tracked.

Top models

#ModelScore
1Claude Opus 4.658
2GPT-554
3GPT-5.253
4Claude Opus 4.552
5Claude Sonnet 447
6Gemini 2.5 Flash41
7Gemini 2.5 Pro38
8GPT-4.137
9Qwen 3 32B31

Interactive version: theaggregate.ai/benchmark?slug=car-bench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.