CAME-Bench - Small: leaderboard

Metric: Answer-set macro F1 (%; 144 questions on six trajectories averaging 23K tokens; full-context answer, gpt-4.1-mini judge). Source: arxiv.org. Saturation forecast: Around December 2026. 5 models tracked.

Top models

#ModelScore
1GPT-5 Mini (Medium)80.4
2GPT-4.1 Mini71.2
3Qwen 3 235B A22B 2507 Instruct70.4
4GPT-4o Mini31.5
5DeepSeek V3.122.8

Interactive version: theaggregate.ai/benchmark?slug=came-bench-small · How It Works · Data refreshed daily, snapshot 2026-09-26.