TriBench-Ko - Hallucination: leaderboard

Metric: Macro-F1 (0-100) on the TriBench-Ko Hallucination risk items (fabricated case citations, statutory provisions or facts): yes/no verification of statements about Korean court decisions, computed over the atomic binary judgments, zero-shot, temperature 0, at most 64 output tokens; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 13 models tracked.

Top models

#ModelScore
1GPT-5.493.3
2Qwen 3.5 9B88.9
3GPT-5.4 Mini88.3
4GPT-4o84.3
5Midm-2.0-Base-Instruct83.3
6Phi-482.6
7Ministral-3-8B-Instruct-251277.2
8Qwen 3 8B70.6
9Gemma 3 12B (IT)68.2
10EXAONE 3.5 7.8B Instruct64.8
11Llama 3.1 8B Instruct37.8

Interactive version: theaggregate.ai/benchmark?slug=tribench-ko-hallucination · How It Works · Data refreshed daily, snapshot 2026-10-07.