AIMultiple - HALC-Bench: leaderboard

Metric: Traps correctly answered 'not mentioned' (%, mean of three haystack positions). Source: aimultiple.com. 11 models tracked.

Top models

#ModelScore
1GPT-5.598.1
2Gemini 3.1 Pro (Preview)96.6
3Kimi K2.694.6
4MiniMax-M2.794.5
5Qwen 3.6 Plus92.2
6GLM-5.187.2
7Claude Opus 4.881.4
8Claude Sonnet 4.673
9Gemini 3.5 Flash67.2
10GPT-5.4 Mini33.3

Interactive version: theaggregate.ai/benchmark?slug=aimultiple-halc-bench · How It Works · Data refreshed daily, snapshot 2026-09-19.