Jian Tang

2 articles on SOTA Papers

Qwen-Audio-3.0-ASR Improves Entity Recall With Hierarchical Hotwords

An instruction-controlled MoE ASR model trained on tens of millions of hours reaches 99.43% recall for priority person names under hotword conditioning.

Sep 12, 20264 min2609.07549

Generative Latents Push Video Compression Below 0.005 Bpp

Group-of-Latents sends sparse anchor latents plus text, then uses a DiT denoiser to synthesize missing video latents at no extra bitrate.

Jul 28, 20265 min2607.19437