LexKairos - Event Relation Discerning: leaderboard

Metric: F1 score (%; 450 items: classify the temporal relation between two events of a judicial case; vanilla setting, the model answers directly without extra reasoning steps (thinking disabled where the model allows it); temperature 0, 8,000 output tokens). Source: arxiv.org. Saturation forecast: Around December 2026. 8 models tracked.

Top models

#ModelScore
1GPT-5.4 (Non-reasoning)90.58
2DeepSeek R187.37
3Gemini 3 Flash (Minimal)82.53
4Qwen 3.5 9B (Non-reasoning)69.94
5DeepSeek V4 Flash (Non-reasoning)54.13
6Llama 3.1 8B28.04

Interactive version: theaggregate.ai/benchmark?slug=lexkairos-event-relation-discerning · How It Works · Data refreshed daily, snapshot 2026-09-29.