Streaming ASR Unifies Transcription and Speaker Attribution
VibeVoice-ASR-Streaming interleaves audio, lookahead, and prior text so one model can produce speaker-attributed transcripts as speech arrives.
Signal processing, systems and control, audio and image processing.
VibeVoice-ASR-Streaming interleaves audio, lookahead, and prior text so one model can produce speaker-attributed transcripts as speech arrives.
Continued V-JEPA-2.1 pretraining on approximately 2,650 hours of surgical video improves partially fine-tuned performance across six robotic-surgery task families.
Radiologist-reviewed masks raise measured nnU-Net Dice scores by 0.143–0.188, exceeding score changes from altering training-data composition.
A 30-hour triple-annotated benchmark separates end-of-turn and interruption decisions, showing that low-latency interruption detection still produces excessive false positives.
An LLM predicts audio latents autoregressively while per-token flow matching generates variable-length scenes, reducing English Seed-TTS WER from 12.15% to 2.79%.
A self-injection-locked 220 GHz autodyne radar forms an intermediate-frequency comb, pairing sub-millimeter separation with 3.4 μm measured ranging accuracy.
A 60 Hz eye-tracking dataset links radiologists’ decision windows to lesions, raising a 3D nnU-Net baseline from 0.6008 to 0.6819 Dice.
A Lego-like 3×4-tile metasurface, guided by a digital twin, redirects 28 GHz links and supports a two-user mmWave demonstration without extra power.
Auditing report-derived MIMIC-CXR labels against radiologist image review found only 1% case capture; a curated-cohort DenseNet121 reached ROC-AUC 0.853.
FullDiT conditions a diffusion transformer on imperfect eight-stream codec plans, lyrics, and captions, improving ViSQOL by 0.77 under synthetic corruption.