Zekai Zhang

2 articles on SOTA Papers

DeepSeek Cuts KV Cache Footprint Fourfold

A causal encoder-decoder MoE pairs cross-layer sparse-attention reuse with FP4 caching to reduce global context memory to 890 bytes per token.

Sep 19, 20264 min2609.19969

Qwen Audio VAE Compresses Audio Without Bottlenecking Training

A 12.5 Hz continuous latent with encoder pruning cuts 64×30 s encoding latency from 1957 ms to 541 ms.

Jul 28, 20265 min2607.11738