TurnBench Exposes Turn-Taking Failures Across Conversation Types
A 30-hour triple-annotated benchmark separates end-of-turn and interruption decisions, showing that low-latency interruption detection still produces excessive false positives.
Signal processing, systems and control, audio and image processing.
A 30-hour triple-annotated benchmark separates end-of-turn and interruption decisions, showing that low-latency interruption detection still produces excessive false positives.
An LLM predicts audio latents autoregressively while per-token flow matching generates variable-length scenes, reducing English Seed-TTS WER from 12.15% to 2.79%.
A self-injection-locked 220 GHz autodyne radar forms an intermediate-frequency comb, pairing sub-millimeter separation with 3.4 μm measured ranging accuracy.
A 60 Hz eye-tracking dataset links radiologists’ decision windows to lesions, raising a 3D nnU-Net baseline from 0.6008 to 0.6819 Dice.
A Lego-like 3×4-tile metasurface, guided by a digital twin, redirects 28 GHz links and supports a two-user mmWave demonstration without extra power.
Auditing report-derived MIMIC-CXR labels against radiologist image review found only 1% case capture; a curated-cohort DenseNet121 reached ROC-AUC 0.853.
FullDiT conditions a diffusion transformer on imperfect eight-stream codec plans, lyrics, and captions, improving ViSQOL by 0.77 under synthetic corruption.
A 260-hour benchmark spanning 21 attacks and five emotions shows conventional detectors can approach chance-level performance under emotional spoofing.
EEG-JEPA predicts multi-depth target representations over structured electrode–time masks, lifting frozen 14-task balanced accuracy from 40.49% to 52.94%.
ECG-guided cross-modal pretraining lets a smartwatch PPG model estimate heart age alone, linking a one-standard-deviation gap increase to 1.72-fold higher hypertension odds.