CAST (Discourse Stress) - Text-Only Hit: leaderboard

Metric: Hit rate (%): items where the predicted stressed word is the intended one; 113 contrastive context pairs (226 items) in which one sentence must stress a different word under each of its two contexts; text-only models name the stressed word zero-shot, non-thinking mode; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 4 models tracked.

Top models

#ModelScore
1Claude Haiku 4.588.1
2Qwen 3 4B (Non-reasoning)57.5
3Qwen 3 1.7B (Non-reasoning)50
4Qwen 3 0.6B (Non-reasoning)25.2

Interactive version: theaggregate.ai/benchmark?slug=cast-discourse-stress-text-only-hit · How It Works · Data refreshed daily, snapshot 2026-10-07.