Skip to content
The Aggregate

Unified LLM rankings, daily

Models What's New Trends Benchmarks Skill Maps Task Explorer
How It Works Guesswork Metrics
The Aggregate
Loading data...

Bongard Problems (Classic) — leaderboard

Metric: Accuracy (% correct). Source: github.com. 4 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)55.4
2Gemini 3 Pro (Preview)50.5
3Gemini 3 Flash (Preview)38.2
4Gemini 2.5 Flash16.2

Interactive version: theaggregate.ai/benchmark?slug=bongard-problems-classic · How It Works · Data refreshed daily, snapshot 2026-07-25.

Built by Mikhail Doroshenko — AI researcher, co-author of Humanity’s Last Exam. Independent project; no affiliation with any AI lab.