Jian Luan

2 articles on SOTA Papers

End-to-End Audio Generation Cuts Speech Errors

An LLM predicts audio latents autoregressively while per-token flow matching generates variable-length scenes, reducing English Seed-TTS WER from 12.15% to 2.79%.

Aug 25, 20264 min2608.11804 Code available

On-Policy Distillation Reduces Cross-Platform GUI Forgetting

Platform-conditioned teacher selection distills desktop and mobile policies into one continual learner, reaching 38.2% OSWorld and 12.0% MobileWorld success.

Jul 12, 20264 min2607.04425