Know2Guess: leaderboard

Metric: Reliability (0-1 scaled to %): the share of test items handled correctly, a correct answer on the answer-expected zones A to C and an abstention on the synthetic-unknown zone D (no hidden gold answer); standard prompt on the 1,080-item test split, greedy decoding with at most 128 or 160 output tokens, no retrieval or tools; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 9 models tracked.

Top models

#ModelScore
1Qwen 3.5 9B (Non-reasoning)57.04
2Mistral Small 3.253.98
3Qwen 3 8B (Non-reasoning)47.22
4Qwen 2.5 3B Instruct38.15
5Llama 3 8B Instruct35.46
6Qwen 2.5 1.5B Instruct15.74

Interactive version: theaggregate.ai/benchmark?slug=know2guess · How It Works · Data refreshed daily, snapshot 2026-09-29.