CzechTopic: leaderboard

Metric: Word-level F1 (%) of the predicted topic spans, on CzechTopic's 1,820 human-annotated (text, topic) pairs (525 historical Czech texts, 363 topics defined by a name and description), zero-shot or two-shot prompting with the authors' span tagging or matching prompts, macro-averaged over topics against the human annotators' spans; higher is better. Source: arxiv.org. Saturation forecast: Around May 2028. 9 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.261.1#105
2GPT-OSS-20B55.4#499
3GPT-5 Mini (2025-08-07)53.4#165
4Llama 3.3 70B Instruct53#520
5Gemma 3 27B50.7#596
6Gemini 3 Pro (Preview)47.4#64
7Gemma 3 4B30.7#1084
8Llama 3.2 3B Instruct29.2#1321
9GPT-5 Nano13.2#415

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=czechtopic · How It Works · Data refreshed daily, snapshot 2026-10-11.