Audio and Speech Processing

eess.AS

Analysis, synthesis, enhancement, and classification of audio and speech.

Sort:

TurnBench Exposes Turn-Taking Failures Across Conversation Types

A 30-hour triple-annotated benchmark separates end-of-turn and interruption decisions, showing that low-latency interruption detection still produces excessive false positives.

Aug 27, 20264 min2608.25218

End-to-End Audio Generation Cuts Speech Errors

An LLM predicts audio latents autoregressively while per-token flow matching generates variable-length scenes, reducing English Seed-TTS WER from 12.15% to 2.79%.

Aug 25, 20264 min2608.11804 Code available

Full-Context Rendering Reduces Music Codec Exposure Bias

FullDiT conditions a diffusion transformer on imperfect eight-stream codec plans, lyrics, and captions, improving ViSQOL by 0.77 under synthetic corruption.

Aug 12, 20264 min2608.08787

Emotionally Expressive Attacks Expose Speech Deepfake Detection Failures

A 260-hour benchmark spanning 21 attacks and five emotions shows conventional detectors can approach chance-level performance under emotional spoofing.

Aug 10, 20264 min2608.05507

Unified Diffusion Audio Improves Structured Scene Control

A DiT conditioned on structured temporal records generates 48 kHz stereo mixtures through 25 Hz VAE latents, raising rich-timeline mIoU to 43.73 from 38.48.

Aug 1, 20265 min2607.27011

Time Markers Ground Audio Answers Across Two Hours

GigaChat Audio interleaves timestamp anchors with continuous audio tokens, reaching 65.2 mIoU on 20–40 minute grounding when trained with 7-second anchors.

Jul 30, 20265 min2607.10387

Audio LLM Feedback Improves Text-to-Audio Instruction Following

ALLM-judged DPO turns event-presence and temporal-order checks into preferences, raising AudioCaps-test joint accuracy to 71.0% from 67.4%.

Jul 28, 20265 min2607.13408

Balanced Pretraining Improves ASR for Central Asian Languages

A Conformer encoder trained on 2M hours uses cluster-level balancing and domain-aware fine-tuning to reduce head-language dominance.

Jul 28, 20264 min2607.10371

GigaSpeechBench Exposes ASR Gaps Beyond Standard Benchmarks

A 680-hour human-annotated benchmark combines multilingual, dialect, accent, domain, and age stress tests across ASR and speech translation.

Jul 28, 20265 min2606.28884

Open AV-LLM Extends Reasoning to Long Videos

AV-Flamingo pairs a 7M-instance audio-visual training set with a three-stage curriculum for multi-event video understanding across 15+ benchmarks.

Jul 28, 20264 min2607.16107