Meng Li

2 articles on SOTA Papers

DeepSeek Cuts KV Cache Footprint Fourfold

A causal encoder-decoder MoE pairs cross-layer sparse-attention reuse with FP4 caching to reduce global context memory to 890 bytes per token.

Sep 19, 20264 min2609.19969

Automated Research Produces Reviewable Papers, With Weak Evidence

FARS runs ideation, planning, experiments, and writing end to end, producing 166 papers with 282 human reviews exposing quality and integrity gaps.

Jul 28, 20265 min2606.31651 Code available