htkgh-polecat (Entity-Filtered History): leaderboard

Metric: Relation prediction accuracy (%; choose the relation of a held-out POLECAT event among 42 candidates, PLOVER event types combined with their modes, given the 100 most recent earlier events that share a primary entity with it (entity filter on, location and context filters off); stratified test set of 5,562 events from 2019 to July 2024; non-thinking models answer in at most 14 tokens, thinking models in up to 16,384). Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1Qwen 3 4B 2507 (Thinking)40.6
2Gemma 3 12B (IT)40
3Qwen 3 8B (Thinking)38.9
4Qwen 3 4B 2507 Instruct36.3
5GPT-OSS-20B (Medium)34.7
6Gemma 3 4B (IT)32.7
7Qwen 3 8B (Non-reasoning)30.9
8Llama 3.1 8B Instruct24.3
9DeepSeek-R1-Distill-Qwen-7B12.2

Interactive version: theaggregate.ai/benchmark?slug=htkgh-polecat-entity-filtered-history · How It Works · Data refreshed daily, snapshot 2026-09-26.