Yichen Jiang

2 articles on SOTA Papers

DeepSeek Cuts KV Cache Footprint Fourfold

A causal encoder-decoder MoE pairs cross-layer sparse-attention reuse with FP4 caching to reduce global context memory to 890 bytes per token.

Sep 19, 20264 min2609.19969

A 35B Agent Matches Trillion-Parameter Models on Long-Horizon Tasks

Long-horizon trajectories, domain teachers, and routed on-policy distillation produce 56.4 on SEAL-0 and 80.6 on IFBench.

Jul 28, 20265 min2606.30616