NYT Connections Extended: leaderboard

Extended version: 940 puzzles with extra trick words added that don't fit any category. Models get one attempt. Significantly harder than the original.

Metric: Score (%). Source: github.com. Status: saturation imminent. 104 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)96.1
2GPT-5.5 (xHigh)96.1
3GPT-5.5 (High)95.2
4GPT-5.4 (xHigh)93.4
5Gemini 3 Pro (Preview)92.3
6Kimi K392.1
7Gemini 3.5 Flash90.8
8Qwen 3.8 Max88.9
9Gemini 3.6 Flash88.5
10GPT-5.4 (High)88.4
11Claude Opus 4.6 (Thinking, High)88.1
12GPT-5.5 (Medium)87.3
13DeepSeek V4 Flash87.2
14Kimi K2.680.2
15Qwen 3.7 Max78.9

Interactive version: theaggregate.ai/benchmark?slug=nyt-connections-extended · How It Works · Data refreshed daily, snapshot 2026-09-05.