SEABED - Long-Form Reasoning (Multiple-Choice): leaderboard
Metric: Accuracy (%; 182 reasoning over whole 7 to 23 minute Thai, Indonesian and Malay conversations items; multiple-choice items scored by exact match against the shuffled gold option; audio at 16 kHz). Source: arxiv.org. Saturation forecast: Around December 2026. 6 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.5 Flash | 80.2 |
| 2 | Gemini 3.6 Flash | 79.1 |
| 3 | Gemini 2.5 Pro | 76.1 |
| 4 | SeaLLMs-Audio-7B | 15.9 |
Interactive version: theaggregate.ai/benchmark?slug=seabed-long-form-reasoning-multiple-choice · How It Works · Data refreshed daily, snapshot 2026-09-26.