TurnBench Exposes Turn-Taking Failures Across Conversation Types
A 30-hour triple-annotated benchmark separates end-of-turn and interruption decisions, showing that low-latency interruption detection still produces excessive false positives.
Analysis, synthesis, enhancement, and classification of audio and speech.
A 30-hour triple-annotated benchmark separates end-of-turn and interruption decisions, showing that low-latency interruption detection still produces excessive false positives.
An LLM predicts audio latents autoregressively while per-token flow matching generates variable-length scenes, reducing English Seed-TTS WER from 12.15% to 2.79%.
FullDiT conditions a diffusion transformer on imperfect eight-stream codec plans, lyrics, and captions, improving ViSQOL by 0.77 under synthetic corruption.
A 260-hour benchmark spanning 21 attacks and five emotions shows conventional detectors can approach chance-level performance under emotional spoofing.
A DiT conditioned on structured temporal records generates 48 kHz stereo mixtures through 25 Hz VAE latents, raising rich-timeline mIoU to 43.73 from 38.48.
GigaChat Audio interleaves timestamp anchors with continuous audio tokens, reaching 65.2 mIoU on 20–40 minute grounding when trained with 7-second anchors.
ALLM-judged DPO turns event-presence and temporal-order checks into preferences, raising AudioCaps-test joint accuracy to 71.0% from 67.4%.
A Conformer encoder trained on 2M hours uses cluster-level balancing and domain-aware fine-tuning to reduce head-language dominance.
A 680-hour human-annotated benchmark combines multilingual, dialect, accent, domain, and age stress tests across ASR and speech translation.
AV-Flamingo pairs a 7M-instance audio-visual training set with a three-stage curriculum for multi-event video understanding across 15+ benchmarks.