Jianwei Yu

4 articles on SOTA Papers

Qwen-Audio-3.0-ASR Improves Entity Recall With Hierarchical Hotwords

An instruction-controlled MoE ASR model trained on tens of millions of hours reaches 99.43% recall for priority person names under hotword conditioning.

Sep 12, 20264 min2609.07549

Streaming ASR Unifies Transcription and Speaker Attribution

VibeVoice-ASR-Streaming interleaves audio, lookahead, and prior text so one model can produce speaker-attributed transcripts as speech arrives.

Sep 4, 20263 min2609.02812

Full-Context Rendering Reduces Music Codec Exposure Bias

FullDiT conditions a diffusion transformer on imperfect eight-stream codec plans, lyrics, and captions, improving ViSQOL by 0.77 under synthetic corruption.

Aug 12, 20264 min2608.08787

GigaSpeechBench Exposes ASR Gaps Beyond Standard Benchmarks

A 680-hour human-annotated benchmark combines multilingual, dialect, accent, domain, and age stress tests across ASR and speech translation.

Jul 28, 20265 min2606.28884