GigaSpeechBench Exposes ASR Gaps Beyond Standard Benchmarks
A 680-hour human-annotated benchmark combines multilingual, dialect, accent, domain, and age stress tests across ASR and speech translation.
Jul 28, 20265 min2606.28884
2 articles on SOTA Papers
A 680-hour human-annotated benchmark combines multilingual, dialect, accent, domain, and age stress tests across ASR and speech translation.
Qwen-Audio-3.0-TTS couples a 12.5 Hz tokenizer with staged LM–FM training, supporting 16 languages and one-pass 3-minute generation.