MOEVE - Topic Extraction: leaderboard
Metric: Mean overall topic-extraction score x 100 (0-100) over two German datasets (KIKC Topics, German Ministry Publications), combining semantic topic match with LLM-judged topic adherence; German prompts, default decoding, open-weight models on vLLM or Ollama at the quantization listed in the paper's model table, others by API; the paper prints only the top 10 of 39 evaluated models; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-4o | 64.1 |
| 2 | Mistral Small 3.1 | 64 |
| 3 | GPT-4o Mini | 63.8 |
| 4 | Llama 3.3 70B | 63.8 |
| 5 | Mistral Large 3 | 62.5 |
| 6 | DeepSeek R1 Distill Qwen 32B | 62.5 |
Interactive version: theaggregate.ai/benchmark?slug=moeve-topic-extraction · How It Works · Data refreshed daily, snapshot 2026-09-29.