Yujie Tu

2 articles on SOTA Papers

Streaming ASR Unifies Transcription and Speaker Attribution

VibeVoice-ASR-Streaming interleaves audio, lookahead, and prior text so one model can produce speaker-attributed transcripts as speech arrives.

Sep 4, 20263 min2609.02812

GigaSpeechBench Exposes ASR Gaps Beyond Standard Benchmarks

A 680-hour human-annotated benchmark combines multilingual, dialect, accent, domain, and age stress tests across ASR and speech translation.

Jul 28, 20265 min2606.28884