SEABED - Long-Form Reasoning (Multiple-Choice): leaderboard

Metric: Accuracy (%; 182 reasoning over whole 7 to 23 minute Thai, Indonesian and Malay conversations items; multiple-choice items scored by exact match against the shuffled gold option; audio at 16 kHz). Source: arxiv.org. Saturation forecast: Around December 2026. 6 models tracked.

Top models

#ModelScore
1Gemini 3.5 Flash80.2
2Gemini 3.6 Flash79.1
3Gemini 2.5 Pro76.1
4SeaLLMs-Audio-7B15.9

Interactive version: theaggregate.ai/benchmark?slug=seabed-long-form-reasoning-multiple-choice · How It Works · Data refreshed daily, snapshot 2026-09-26.