Confetti (Voice): leaderboard
Metric: AST soft accuracy (%; exact match on function name and non-string arguments, AlignScore on string arguments; 313 spoken Confetti queries, mean over six TTS voice configurations). Source: arxiv.org. Saturation forecast: Around December 2026. 7 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Qwen3 Omni 30B A3B Instruct | 60.4 |
| 2 | Phi-4 Multimodal Instruct | 23.3 |
Interactive version: theaggregate.ai/benchmark?slug=confetti-voice · How It Works · Data refreshed daily, snapshot 2026-09-26.