FOA Tokens Improve Spatial Audio Reasoning Without Replacing Audio Encoders
A parallel FOA encoder adds spatial latents to Omni LLMs, raising Qwen3-Omni's MMAU-Pro spatial score by 9.50 points while retaining general-audio performance.
A parallel FOA encoder adds spatial latents to Omni LLMs, raising Qwen3-Omni's MMAU-Pro spatial score by 9.50 points while retaining general-audio performance.
An incentive-compatibility-in-the-large characterization turns a high-dimensional reporting problem into a one-dimensional acceptance rule with deliberate surplus burning.
Rank-2 log-unit lattices encode Galois structure and, for totally imaginary non-CM D6 sextics, determine the underlying field.
A curated 830-hour dialogue dataset couples style-conditioned instructions with synthesized speech, raising multi-dimensional control scores without reported losses in conversational capability.
Structural exclusions yield a finite path-width criterion, while universal-host constructions separate line-width bounds from path-width guarantees.
A 300-task procedural suite replaces preference-only judging with task-specific scorers, raising matched reinforcement-learning performance from 0.509 to 0.548.
VibeVoice-ASR-Streaming interleaves audio, lookahead, and prior text so one model can produce speaker-attributed transcripts as speech arrives.
Gram-matrix positivity converts the boundary identity into finite-dimensional linear algebra, yielding an exact two-parameter admissibility region and a complete dimension classification.