SocialMemBench (Full Context, Small-Tier Subset): leaderboard

Metric: Mean score (%; question-weighted over 77 questions of 10 small-tier social networks, whole conversation corpus in context; multiple choice by exact letter match, open-ended by a GPT-4o-mini judge in 0-1). Source: arxiv.org. Saturation forecast: Around December 2026. 2 models tracked.

Top models

#ModelScore
1Gemini 2.5 Flash71.9
2GPT-4o Mini64.4

Interactive version: theaggregate.ai/benchmark?slug=socialmembench-full-context-small-tier-subset · How It Works · Data refreshed daily, snapshot 2026-09-26.