htkgh-polecat (Entity-Filtered History): leaderboard
Metric: Relation prediction accuracy (%; choose the relation of a held-out POLECAT event among 42 candidates, PLOVER event types combined with their modes, given the 100 most recent earlier events that share a primary entity with it (entity filter on, location and context filters off); stratified test set of 5,562 events from 2019 to July 2024; non-thinking models answer in at most 14 tokens, thinking models in up to 16,384). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen 3 4B 2507 (Thinking) | 40.6 |
| 2 | Gemma 3 12B (IT) | 40 |
| 3 | Qwen 3 8B (Thinking) | 38.9 |
| 4 | Qwen 3 4B 2507 Instruct | 36.3 |
| 5 | GPT-OSS-20B (Medium) | 34.7 |
| 6 | Gemma 3 4B (IT) | 32.7 |
| 7 | Qwen 3 8B (Non-reasoning) | 30.9 |
| 8 | Llama 3.1 8B Instruct | 24.3 |
| 9 | DeepSeek-R1-Distill-Qwen-7B | 12.2 |
Interactive version: theaggregate.ai/benchmark?slug=htkgh-polecat-entity-filtered-history · How It Works · Data refreshed daily, snapshot 2026-09-26.