Jinpeng Wang

2 articles on SOTA Papers

DeepSeek Cuts KV Cache Footprint Fourfold

A causal encoder-decoder MoE pairs cross-layer sparse-attention reuse with FP4 caching to reduce global context memory to 890 bytes per token.

Sep 19, 20264 min2609.19969

On-Policy Distillation Reduces Cross-Platform GUI Forgetting

Platform-conditioned teacher selection distills desktop and mobile policies into one continual learner, reaching 38.2% OSWorld and 12.0% MobileWorld success.

Jul 12, 20264 min2607.04425