SpeechEditBench - Compositional (Two Edits): leaderboard

Metric: Joint success (%): the requested edit is applied (task-specific automatic check) and the linguistic content is preserved (Whisper large-v3 or Paraformer transcript within 10% WER or CER of the expected text); source speech plus a natural-language instruction, English and Mandarin samples, each model through its released editing or conversation interface; two-edit instructions without a speaker component (content, emotion, prosody and acoustic pairs, 200 samples), every requested edit must succeed and the content must be preserved; higher is better. Source: arxiv.org. Saturation forecast: Around August 2028. 8 models tracked.

Top models

#ModelScore
1Qwen3 Omni 30B A3B Instruct10

Interactive version: theaggregate.ai/benchmark?slug=speecheditbench-compositional-two-edits · How It Works · Data refreshed daily, snapshot 2026-09-29.