SpeechEditBench - Style: leaderboard
Metric: Joint success (%): the requested edit is applied (task-specific automatic check) and the linguistic content is preserved (Whisper large-v3 or Paraformer transcript within 10% WER or CER of the expected text); source speech plus a natural-language instruction, English and Mandarin samples, each model through its released editing or conversation interface; speaking-style editing (six styles such as broadcast, intimate and storytelling, judged by a gemini-2.5-pro audio judge), 600 samples; higher is better. Source: arxiv.org. Saturation forecast: Around July 2027. 8 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen3 Omni 30B A3B Instruct | 24.17 |
Interactive version: theaggregate.ai/benchmark?slug=speecheditbench-style · How It Works · Data refreshed daily, snapshot 2026-09-29.