Condesion-Bench - Contextual Conditions: leaderboard

Metric: Satisfaction rate (%) of contextual conditions (feasibility tied to the scenario context), of the single action an LLM generates zero-shot (temperature 0, JSON output) on the 251 Condesion-Bench instances (one per trading day, October 2024 to September 2025: allocate a budget over 15 S&P 500 stocks under variable, contextual and allocation conditions); higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 16 models tracked.

Top models

#ModelScore
1GPT-5 Mini97.21
2GPT-596.81
3O4 Mini (2025-04-16)96.41
4O3 (2025-04-16)96.41
5Claude 3.5 Sonnet (20241022)93.63
6GPT-OSS-120B86.06
7Claude 3.5 Haiku85.66
8GPT-OSS-20B84.46
9GPT-4.179.68
10Gemini 2.0 Flash (001)77.29
11GPT-4.1 Mini76.1
12Llama 3.3 70B Instruct72.51
13Mistral Large 2 (Jul)71.71
14Mistral Small 3.266.53
15Gemini 2.0 Flash Lite (001)60.16

Interactive version: theaggregate.ai/benchmark?slug=condesion-bench-contextual-conditions · How It Works · Data refreshed daily, snapshot 2026-10-07.