SEABED - Long-Form Reasoning (Open-Ended): leaderboard
Metric: Accuracy (%; 182 reasoning over whole 7 to 23 minute Thai, Indonesian and Malay conversations items; open-ended answers judged by Grok 4.5, binary correct or incorrect (graded 0 to 1 for long-form); audio at 16 kHz). Source: arxiv.org. Saturation forecast: Around February 2027. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.6 Flash | 68.3 |
| 2 | Gemini 3.5 Flash | 66.7 |
| 3 | Gemini 2.5 Pro | 64.6 |
| 4 | SeaLLMs-Audio-7B | 11.3 |
Interactive version: theaggregate.ai/benchmark?slug=seabed-long-form-reasoning-open-ended · How It Works · Data refreshed daily, snapshot 2026-09-26.