Computer ScienceDeepSeek Cuts KV Cache Footprint FourfoldA causal encoder-decoder MoE pairs cross-layer sparse-attention reuse with FP4 caching to reduce global context memory to 890 bytes per token.cs.CLSep 19, 20264 min2609.19969