Jingbei Li

2 articles on SOTA Papers

Qwen-Audio-3.0-ASR Improves Entity Recall With Hierarchical Hotwords

An instruction-controlled MoE ASR model trained on tens of millions of hours reaches 99.43% recall for priority person names under hotword conditioning.

Sep 12, 20264 min2609.07549

Unified Diffusion Audio Improves Structured Scene Control

A DiT conditioned on structured temporal records generates 48 kHz stereo mixtures through 25 Hz VAE latents, raising rich-timeline mIoU to 43.73 from 38.48.

Aug 1, 20265 min2607.27011