Bongard Problems (Classic): leaderboard

Metric: Accuracy (% correct). Source: github.com. 15 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)55.4
2Gemini 3 Pro (Preview)50.5
3Gemini 3 Flash (Preview)38.2
4Gemini 2.5 Flash16.2
5Mistral Medium 3.111.5
6Ministral 3 14B9.8
7Mistral Medium 39.3
8Ministral 3 8B8.6
9Mistral Small 3.28.3
10Mistral Medium 3.56.9
11Magistral Small 1.26.6
12Magistral Medium 1.26.6
13Mistral Small 45.6
14Mistral Large 35.1

Interactive version: theaggregate.ai/benchmark?slug=bongard-problems-classic · How It Works · Data refreshed daily, snapshot 2026-09-08.