Yu Zhang

5 articles on SOTA Papers

DeepSeek Cuts KV Cache Footprint Fourfold

A causal encoder-decoder MoE pairs cross-layer sparse-attention reuse with FP4 caching to reduce global context memory to 890 bytes per token.

Sep 19, 20264 min2609.19969

FOA Tokens Improve Spatial Audio Reasoning Without Replacing Audio Encoders

A parallel FOA encoder adds spatial latents to Omni LLMs, raising Qwen3-Omni's MMAU-Pro spatial score by 9.50 points while retaining general-audio performance.

Sep 9, 20264 min2606.10738 Code available

Agent Swarm Cuts Complex Task Latency Up To 4.5×

Kimi K2.5 couples joint text-vision training with dynamically scheduled parallel subagents, improving agentic search scores while reducing time to target quality.

Aug 13, 20264 min2602.02276

Kimi K3 Brings Open Weights Closer to Frontier Agents

A 2.8T-parameter MoE combines Delta Attention, 16-of-896 expert routing, and agentic reinforcement learning to approach leading proprietary systems at lower task cost.

Aug 10, 20264 min2607.24653

Mini Activations Narrow Frontier Model Gaps

A sparse 229.9B-parameter MoE activates 9.8B parameters per token and pairs agent-native data with RL for coding, cowork, and reasoning tasks.

Jul 31, 20265 min2605.26494