Lu Chen

4 articles on SOTA Papers

DeepSeek Cuts KV Cache Footprint Fourfold

A causal encoder-decoder MoE pairs cross-layer sparse-attention reuse with FP4 caching to reduce global context memory to 890 bytes per token.

Sep 19, 20264 min2609.19969

MinerU-Chem Reports 93.02% Exact Molecular Recognition

A chemistry-focused extension to MinerU converts document regions into molecule and reaction records, outperforming the evaluated GPT-5.6-Sol comparison on a SMILES benchmark subset.

Sep 14, 20262 min2608.03525

Mini Activations Narrow Frontier Model Gaps

A sparse 229.9B-parameter MoE activates 9.8B parameters per token and pairs agent-native data with RL for coding, cowork, and reasoning tasks.

Jul 31, 20265 min2605.26494

Evidence-Grounded Agents Improve Musculoskeletal Care Pathways

OrthoPilot combines hospital data retrieval with external medical knowledge, improving physician management success by 10.6% in 1,870 prospective complex cases.

Jul 28, 20264 min2607.12527