Qwen-Audio-3.0-ASR Improves Entity Recall With Hierarchical Hotwords
An instruction-controlled MoE ASR model trained on tens of millions of hours reaches 99.43% recall for priority person names under hotword conditioning.
5 articles on SOTA Papers
An instruction-controlled MoE ASR model trained on tens of millions of hours reaches 99.43% recall for priority person names under hotword conditioning.
FullDiT conditions a diffusion transformer on imperfect eight-stream codec plans, lyrics, and captions, improving ViSQOL by 0.77 under synthetic corruption.
A DiT conditioned on structured temporal records generates 48 kHz stereo mixtures through 25 Hz VAE latents, raising rich-timeline mIoU to 43.73 from 38.48.
A 680-hour human-annotated benchmark combines multilingual, dialect, accent, domain, and age stress tests across ASR and speech translation.
Qwen-Audio-3.0-TTS couples a 12.5 Hz tokenizer with staged LM–FM training, supporting 16 languages and one-pass 3-minute generation.