DeepSeek Cuts KV Cache Footprint Fourfold
A causal encoder-decoder MoE pairs cross-layer sparse-attention reuse with FP4 caching to reduce global context memory to 890 bytes per token.
Sep 19, 20264 min2609.19969
3 articles on SOTA Papers
A causal encoder-decoder MoE pairs cross-layer sparse-attention reuse with FP4 caching to reduce global context memory to 890 bytes per token.
A parallel FOA encoder adds spatial latents to Omni LLMs, raising Qwen3-Omni's MMAU-Pro spatial score by 9.50 points while retaining general-audio performance.
A sparse 229.9B-parameter MoE activates 9.8B parameters per token and pairs agent-native data with RL for coding, cowork, and reasoning tasks.