TriBench-Ko - Prompt Sensitivity: leaderboard

Metric: Macro-F1 (0-100) on the TriBench-Ko Prompt Sensitivity risk items (the same prompt issued five times, scored strictly over the repetitions): yes/no verification of statements about Korean court decisions, computed over the atomic binary judgments, zero-shot, temperature 0, at most 64 output tokens; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 13 models tracked.

Top models

#ModelScore
1GPT-5.492.3
2Qwen 3.5 9B87.1
3GPT-5.4 Mini87.1
4GPT-4o79.7
5Midm-2.0-Base-Instruct78.2
6Phi-475.7
7Ministral-3-8B-Instruct-251271.1
8Gemma 3 12B (IT)60.5
9Qwen 3 8B57.5
10EXAONE 3.5 7.8B Instruct57.2
11Llama 3.1 8B Instruct35

Interactive version: theaggregate.ai/benchmark?slug=tribench-ko-prompt-sensitivity · How It Works · Data refreshed daily, snapshot 2026-10-07.