CLIR-Bench - Intervention Response: leaderboard

Metric: Accuracy (%) on intervention response, 4-option multiple-choice questions over serialized irregular ICU time series (Full-TS input), chance 25; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 15 models tracked.

Top models

#ModelScore
1Kimi K2.693.33
2DeepSeek V4 Flash91.67
3Gemma 3 27B (IT)77.83
4Gemma 4 12B (IT)60.33
5Gemini 2.5 Flash60
6GPT-5.4 Mini51.67
7Qwen 3.5 9B27.17
8Qwen 3.6 27B23.83
9Qwen 3.5 4B22.5

Interactive version: theaggregate.ai/benchmark?slug=clir-bench-intervention-response · How It Works · Data refreshed daily, snapshot 2026-09-29.