Condesion-Bench (April-September 2025 Subset): leaderboard
Metric: Decision Satisfaction Rate (%, from a 0-1 share times 100) on the 126 Condesion-Bench instances from April to September 2025, sampled to postdate the knowledge cutoffs of recent models (zero-shot, temperature 0, JSON output); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Grok 4 | 89.68 |
| 2 | Gemini 2.5 Pro (Medium) | 88.8 |
| 3 | Grok 4 Fast (Reasoning) | 84 |
| 4 | Claude Sonnet 4.5 (Thinking) | 59.68 |
| 5 | Claude Haiku 4.5 (Thinking) | 56.8 |
| 6 | Gemini 2.5 Flash (Medium) | 43.2 |
| 7 | Nova Pro | 36.8 |
Interactive version: theaggregate.ai/benchmark?slug=condesion-bench-april-september-2025-subset · How It Works · Data refreshed daily, snapshot 2026-10-07.