CAME-Bench - Medium: leaderboard

Metric: Answer-set macro F1 (%; 168 questions on six trajectories averaging 137K tokens; full context, left-truncated to the model's window, gpt-4.1-mini judge). Source: arxiv.org. Saturation forecast: Around December 2026. 5 models tracked.

Top models

#ModelScore
1GPT-5 Mini (Medium)56.6
2GPT-4.1 Mini36.2
3Qwen 3 235B A22B 2507 Instruct15.2
4GPT-4o Mini7.5
5DeepSeek V3.11

Interactive version: theaggregate.ai/benchmark?slug=came-bench-medium · How It Works · Data refreshed daily, snapshot 2026-09-26.