Yang Zhang

2 articles on SOTA Papers

DeepSeek Cuts KV Cache Footprint Fourfold

A causal encoder-decoder MoE pairs cross-layer sparse-attention reuse with FP4 caching to reduce global context memory to 890 bytes per token.

Sep 19, 20264 min2609.19969

FLARE Extends Laboratory Access to Multiple X-Line Reconnection

A larger, modular reconnection experiment combines segmented coils, independent ohmic heating, and upgraded diagnostics to target $S\sim10^5$ and $\lambda\sim10^3$.

Aug 26, 20265 min2608.17332