CAST (Discourse Stress) - Text-Only Hit: leaderboard
Metric: Hit rate (%): items where the predicted stressed word is the intended one; 113 contrastive context pairs (226 items) in which one sentence must stress a different word under each of its two contexts; text-only models name the stressed word zero-shot, non-thinking mode; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 4 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Haiku 4.5 | 88.1 |
| 2 | Qwen 3 4B (Non-reasoning) | 57.5 |
| 3 | Qwen 3 1.7B (Non-reasoning) | 50 |
| 4 | Qwen 3 0.6B (Non-reasoning) | 25.2 |
Interactive version: theaggregate.ai/benchmark?slug=cast-discourse-stress-text-only-hit · How It Works · Data refreshed daily, snapshot 2026-10-07.