Jiajun Zhang

3 articles on SOTA Papers

Streaming ASR Unifies Transcription and Speaker Attribution

VibeVoice-ASR-Streaming interleaves audio, lookahead, and prior text so one model can produce speaker-attributed transcripts as speech arrives.

Sep 4, 20263 min2609.02812

Qwen-CUA Reaches 86.2 on Verified Computer Use

A 397B-A17B mixture-of-experts agent learns screenshot-only keyboard and mouse control from verifiable interactive trajectories, improving OSWorld-Verified performance to 86.2.

Aug 8, 20264 min2608.02352

GigaSpeechBench Exposes ASR Gaps Beyond Standard Benchmarks

A 680-hour human-annotated benchmark combines multilingual, dialect, accent, domain, and age stress tests across ASR and speech translation.

Jul 28, 20265 min2606.28884