Hebrew LLM - Winograd (0-shot): leaderboard

Metric: Accuracy (%). Source: huggingface.co. 34 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)96.04
2Gemini 3 Flash (Preview)94.6
3GLM-5.3 Flash93.88
4Qwen 3.5 122B A10B93.17
5GLM-5.191.37
6GPT-5.191.01
7Gemini 2.5 Pro90.65
8GPT-4o89.21
9Claude Sonnet 4.588.85
10GPT-5 Mini88.49
11Qwen 3.5 397B A17B86.69
12Gemini 2.5 Flash85.97
13Gemma 4 31B (IT)85.97
14c4ai-command-a-03-202584.89
15Llama 3.3 70B Instruct83.45

Interactive version: theaggregate.ai/benchmark?slug=hebrew-llm-winograd-0-shot · How It Works · Data refreshed daily, snapshot 2026-09-05.