SEABED - Long-Form Reasoning (Open-Ended): leaderboard

Metric: Accuracy (%; 182 reasoning over whole 7 to 23 minute Thai, Indonesian and Malay conversations items; open-ended answers judged by Grok 4.5, binary correct or incorrect (graded 0 to 1 for long-form); audio at 16 kHz). Source: arxiv.org. Saturation forecast: Around February 2027. 6 models tracked.

Top models

#ModelScore
1Gemini 3.6 Flash68.3
2Gemini 3.5 Flash66.7
3Gemini 2.5 Pro64.6
4SeaLLMs-Audio-7B11.3

Interactive version: theaggregate.ai/benchmark?slug=seabed-long-form-reasoning-open-ended · How It Works · Data refreshed daily, snapshot 2026-09-26.