CzechTopic - Text-Level F1: leaderboard

Metric: Text-level F1 (%) for deciding whether the topic is present in the text, on CzechTopic's 1,820 human-annotated (text, topic) pairs (525 historical Czech texts, 363 topics defined by a name and description), zero-shot or two-shot prompting with the authors' span tagging or matching prompts, macro-averaged over topics against the human annotators' spans; higher is better. Source: arxiv.org. Saturation forecast: Around November 2027. 9 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.280.6#105
2Llama 3.3 70B Instruct75.2#520
3GPT-OSS-20B74.7#499
4Gemma 3 27B70#596
5GPT-5 Mini (2025-08-07)67.8#165
6Gemini 3 Pro (Preview)62.4#64
7Gemma 3 4B59.9#1084
8Llama 3.2 3B Instruct59.7#1321
9GPT-5 Nano18.4#415

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=czechtopic-text-level-f1 · How It Works · Data refreshed daily, snapshot 2026-10-11.