Unified Diffusion Audio Improves Structured Scene Control
A DiT conditioned on structured temporal records generates 48 kHz stereo mixtures through 25 Hz VAE latents, raising rich-timeline mIoU to 43.73 from 38.48.
Aug 1, 20265 min2607.27011
2 articles on SOTA Papers
A DiT conditioned on structured temporal records generates 48 kHz stereo mixtures through 25 Hz VAE latents, raising rich-timeline mIoU to 43.73 from 38.48.
Qwen-Audio-3.0-TTS couples a 12.5 Hz tokenizer with staged LM–FM training, supporting 16 languages and one-pass 3-minute generation.