CzechTopic - Word-Level IoU: leaderboard

Metric: Word-level intersection over union (%) of the predicted and annotated topic spans, on CzechTopic's 1,820 human-annotated (text, topic) pairs (525 historical Czech texts, 363 topics defined by a name and description), zero-shot or two-shot prompting with the authors' span tagging or matching prompts, macro-averaged over topics against the human annotators' spans; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.248.7#105
2GPT-OSS-20B42.1#499
3GPT-5 Mini (2025-08-07)40.9#165
4Llama 3.3 70B Instruct39.2#520
5Gemma 3 27B36.8#596
6Gemini 3 Pro (Preview)36.6#64
7Gemma 3 4B19.3#1084
8Llama 3.2 3B Instruct18.6#1321
9GPT-5 Nano10.3#415

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=czechtopic-word-level-iou · How It Works · Data refreshed daily, snapshot 2026-10-11.