Grip on LLMs - HonestCityBench - Non-Text Modality: leaderboard

Metric: Appropriate acknowledgement rate (%; LLM judge). Source: arxiv.org. 31 models tracked.

Top models

#ModelScore
1GPT-4o51
2Gemma 3 12B (IT)38
3Phi-4 Mini Instruct37
4GPT-4o Mini35
5Gemma 3 27B (IT)31
6Qwen 3 32B24
7OLMo-2-1124-7B-Instruct24
8DeepSeek R1 Distill Qwen 32B20
9Mistral Small 320
10GEITje-7B-ultra18
11Llama 3.1 8B Instruct15
12command-r7B-12-202415
13aya-expanse-32B14
14Mistral Medium 313
15GPT-OSS-20B10

Interactive version: theaggregate.ai/benchmark?slug=grip-on-llms-honestcitybench-non-text-modality · How It Works · Data refreshed daily, snapshot 2026-09-19.