Hui Wang

5 articles on SOTA Papers

Pruned CTC Cuts 180K-Vocabulary ASR Training Memory

Restricting CTC alignment states to batch target tokens preserves full-vocabulary loss and gradients while cutting full-step memory 5.1× at 180K tokens.

Oct 4, 20264 min2609.33645

VoiceChat Preserves Turn Taking With Native Tool Calls

A streaming speech stack separates response text, function calls, transcription, and synthesis, reaching 82.5% tool-selection F1 while maintaining interruption handling.

Oct 4, 20265 min2609.21967

Mobius Separates Knowledge From Reasoning for Faster Inference

A shared feed-forward memory and iterative attention reasoners retain comparable downstream scores while reducing continual-pretrained 35B-model inference time by nearly 4×.

Aug 25, 20264 min2608.14290

Agent Swarm Cuts Complex Task Latency Up To 4.5×

Kimi K2.5 couples joint text-vision training with dynamically scheduled parallel subagents, improving agentic search scores while reducing time to target quality.

Aug 13, 20264 min2602.02276

Kimi K3 Brings Open Weights Closer to Frontier Agents

A 2.8T-parameter MoE combines Delta Attention, 16-of-896 expert routing, and agentic reinforcement learning to approach leading proprietary systems at lower task cost.

Aug 10, 20264 min2607.24653