DeepSeek Cuts KV Cache Footprint Fourfold

A causal encoder-decoder MoE pairs cross-layer sparse-attention reuse with FP4 caching to reduce global context memory to 890 bytes per token.

Editorial Desk·September 19, 2026·4 min readmoderate

Underlying Paper

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

DeepSeek-AI:Anyi XuB. LiBangcai LinBing XueBingCheng XianBingzheng XuBochao WuBowei ZhangBoyi DengC. C. YuChao JinChaofan LinChen DongChenbing WangChenfan FengChengda LuChenggang ZhaoChengqi DengChengyuan ZhangChenhao XuChenqi ZhaoChenze ShaoChuhao WangChuqi ZhangDamai DaiDejian YangDeli ChenDi HuangDi WuDonghao LiErhang LiEric FuF. ZhouFangwei ZhouFangyun LinFangzhou YuanFeiyu XiaFucong DaiGuangbo HaoGuanglin LiGuanting ChenGuoai CaoGuofan FanGuolai MengGuowei LiHaichuan ZhangHaiyang MaHaiyang ShenHan LiHan YuHan ZhangHangyuan DengHanwei XuHanxiang XuHanxun ZhongHao GuoHao JiangHao LiHao QinHaodong WenHaofen LiangHaofeng HuangHaohua LiuHaoling ZhangHaoming LuoHaoran YangHaotian XuHaotian YuanHaoting HuangHaowen LuoHaoyang CaiHaoyu ChenHaozhe JiHengran ZhangHengrui WangHengxu WuHonghui DingHongxuan TangHuadong WangHuanqi CaoHuazuo GaoHui QuHui ZengJ. YangJ. H. JinJ. H. ZhangJ. X. ZouJia YuJiahui ZhouJiajun ChenJialiang HuangJialin ZhaoJiamin TangJian ZhouJianan TongJianwen LiJiaqi ZhuJiarui WangJiasheng YeJiashi LiJiaxin XuJiaying DingJibai LuJiewen HuJin YanJincheng ZhaiJingchang ChenJingcheng HuJingli ZhouJingsheng XuJingting XiangJingyan YunJingyang YuanJingyuan ChengJinhua ZhuJinpeng WangJinyi ChenJinyi HuJiping YuJueliang GuoJunbo PeiJunbo SunJunguang JiangJunjie QiuJunkang ZhouJunqi LiuJunren LiJunxian LiJunxiao SongJunyi GuoKai DongKaifeng ChenKaige GaoKang GuanKangdong YuanKe HongKe XuKefan ZhaoKexin JiKexin ZhangKexing ZhouKuai YuLan ZhangLean WangLecong ZhangLei WangLetian GaoLiang ZhaoLiansheng XuLihua GuoLingxiao LuoLingyue FuLitao DengLitong WangLiyue ZhangLonghao ChenLu ChenLuotian HuangLuyao MaLuyao WangM. S. DiMax MeiMenghao YeMiao CuiMingchuan ZhangMinghua ZhangMinghui TangMingjing ZhangMingqi WeiMingshu ChenMingxing LiuMingxu ZhouMingyu XuMingyu YangMingze WangMuyang ChenNi ShentuNing WangNiufang NingPanpan HuangPeixin CongPeiyi WangPeiyuan XinPengfei RenPengfei YanPengle ZhangQi KangQi TangQiancheng WangQiang LiQihao ZhuQingyang LiQinyu ChenQiushi DuQizhou GuoRongxian XuRui DingRui HuRui TianRui YuRuidong ZhuRuifan XuRuihan YangRuihang XiaRuijie LuRuilin GengRuipeng HongRuiqi GeRuisong ZhangRuize SunRuizhe PanRunji WangRunqian ChenRunxin XuRuohong TianRuomeng ShenRuoyu ZhangRyan X.S. H. LiuShanghao LuShangyan ZhouShanhuang ChenShaofei CaiShaoheng NieShaoyuan ChenShengding HuShengkai LinShengwen RanShengyu LiuShengyuan JiaShi BaiShi FengShicheng XuShichun LiuShiqiang HuShirong MaShiyu WangShiyuan FengShufan GongShuhan LinShuiping YuShunfeng ZhouShuo YangShuomeng WangShuting GuoShuting PanShuying YuSinuo CaoSiyi LinSizhe ChenSongyang ChenSongyang ZhouTao NiTao YunTian JinTian PeiTian YeTianle LinTianran JiTianyi CuiTianyuan YueTingting YuTongrui XiongWangding ZengWei LiuWei ZhangWeibin XuWeihao ZengWeilin ZhaoWen LiuWenfeng LiangWenjie PangWenjing LuoWenjing YaoWenjun GaoWenkai ShaoWenkai YangWenli ZhangWenlu WangWenlve HuangWenqian YanWentao ZhangXi GaoXiang HeXiang LiXiangli LiXiangwen WangXiangying ZhangXiankui WeiXiao BiXiaodong LiuXiaohan WangXiaojian QuXiaokang ChenXiaokang ZhangXiaotao NieXiaoyao ZouXiaoyuan LiXicheng GuoXieting ChuXin ChengXin LiuXin XieXinbo XuXingchao LiuXingchen LiuXingkai YuXingyou LiXintong YaoXinyang ChenXinyong JiangXinyu YangXinyu YangXu ChenXuanyu WangXubei ZhongXuecheng SuXuejie LiuXuheng LinXujie FanXuncheng ZhaoXuwei FuY. C. YanY. H. JiangY. T. WuY. W. M.Y. Z. WangYafei GaoYang YangYang ZhangYanru MaYanwen HuangYao LiYao LiYao MengYao ZhaoYaofeng SunYaohui WangYaoyang YeYehang YinYexinrui WuYi QianYi TaoYi YuYichao ZhangYichen JiangYicheng WangYifan DingYifan ShiYifeng PengYifeng ZhaiYijia WuYiliang XiongYilun WangYing HeYing ZhouYingjia LuoYinmin ZhongYiping WangYisong WangYixiang ZhangYixiao ChenYixuan TanYixuan WeiYiyang MaYiyao YangYiyuan LiuYizai CaiYizhen WeiYizhi WangYonglun YangYongqi ZhuoYongqiang GuoYongtong WuYu WuYu ZhangYuan BianYuan ChengYuan OuYuan SunYuanfan XuYuanhang SunYuanhao LiYuchen LiuYuchen YaoYudong HanYuduan WangYuhan WuYuhao MengYuheng ZouYuKun LiYunchuan WangYunfan XiaoYunfan XiongYupeng ChenYuqian CaoYuqian WangYuqing ChenYushun ZhangYutong LinYuwei XiaoYuxian GuYuxiang ChenYuxiang HuangYuxiang LuoYuxiang YouYuxin ChenYuxin XiangYuxuan LiuYuxuan ZhouYuyang ZhouYuzhe GuoYuzhen HuangYuzhuo BaiZ. Y. Z.Zanlin NiZehao WangZehua ZhaoZehui RenZejun ZhaoZhangli ShaZhanying WangZhaochen ZhangZhaoshuai DuZhe FuZhean XuZhenda XieZheng LiuZhengyan ZhangZhenhua DongZhewen HaoZhibang WangZhibin GouZhicheng MaZhihao LiZhihong ShaoZhihuan HuangZhijie LiZhirui LuZhixian HuangZhixuan ChenZhixuan ChenZhixuan PanZhiyu WuZhizhou RenZhu HeZhuoshu LiZhuping ZhangZian XuZihao WangZihui GuZijia ZhuZili ZhangZilin LiZilong HouZilong LyuZiqiao WangZiwei XieZiya ZhangZiyi GaoZizheng PanZonglin LiZongqing YaoZui ChenZuofan WuChenchen LingChengyu HouChong ChenD. LiDi QiDongjie JiFang WeiFanyi XiaFei XieFeiyi TanHailong GuoHaiyan ZhaiHui ZhouHuihui TanHuijie LiJia LuoJia SongJialu CaiJian LiangJiangting ZhouJiaqi GaoJiayi ShaoJie ChenJieyu YangJin ChenJingde ZhangJingzi ZhouJinqian WangJinyang LiuJinZhao SunJunhua LingJunmin ZhengKaicheng YangKe XuLe SuLeyi XiaLiangfeng DingLin ZhuoLinwang MaLinyan ZhuLiyu CaiLuqi YaoM. K. ZhangMeng LiMiao LinMiaojun WangMin ZhangMingming LiMingming WangMingze YinMinmin HanNan CaoNing WangNingxin MaPanpan WangPeihan LinPeng SunPeng ZhangQian YingQiang XiangQiao WangQingmiao MaoQiwei JiangRongli JinRuyi ChenSha TaoShangmian SunShaoqing WuShichao ZouSi LeiTianyang ZhangTianyu SunTingting YinW. L. XiaoWei AnWei LiWei WangWeiwei LinWenqing HouX. LinXiangfei MengXianzhu HuangXiao PengXiaoqian LiXiaoting ZhangXiaowen SunXiaoxiang WangXiaoyu YeXinrou ZhangXinyu ZhangXue CaoXueyin ChenYanan ZhouYanhong XuYao XiaYao XuYi ShaoYihong ZhangYiling MaYing TangYining LouYiru ChenYishi PiaoYixuan ChenYong XiongYuchen XuanYuehan YangYuer XuYukun ZhaYunxian MaYuping LinYuting YanYutong XieYuwen ShengYuxuan ZhuZekai ZhangZhe JuZhenzhen LinZheren GaoZheyang SunZhigang YanZhongyu WuZi WangZihua QuZiling YanZiyi Wan

The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Together, these compute, storage, and bandwidth demands constitute the primary bottleneck to further lowering deployment costs. To address this challenge, we introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. With its Causal Encoder-Decoder (CED) architecture, the model activates 16B parameters per token during decode but only 8B parameters during prefill, substantially improving cost efficiency for agentic workloads. To push the limits of KV cache compression, DeepSeek-V4.1-Flash combines cross-layer KV cache reuse in Compressed Sparse Attention 2 (CSA2) with FP4 KV caching. These designs reduce its global KV cache footprint (always in HBM) to 890 bytes per token, roughly 1/4 of the corresponding footprint of DeepSeek-V4-Flash. Further, through a dedicated deployment optimization known as SWA Bounded Replay, DeepSeek-V4.1-Flash reduces its persistent KV cache footprint (always on SSD or in host memory) to roughly 1/8 of that of DeepSeek-V4-Flash. Despite its much smaller KV cache footprint, the model delivers substantially better performance than the baseline. In addition, we streamline the DeepSeek-V4 architecture and introduce several efficient architectural extensions. We pretrain DeepSeek-V4.1-Flash on a multimodal corpus comprising 45T tokens and conduct comprehensive post-training, yielding strong performance across diverse text-based and multimodal agentic scenarios. Model checkpoints are available at https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash.

arXiv:2609.19969Submitted: Sep 18, 2026v1

Long-horizon agents turn context handling into a deployment constraint rather than a secondary systems detail. Prefill must process growing histories, while key-value caches occupy accelerator memory, host memory, or SSD and must be moved across those tiers during inference. DeepSeek-V4.1-Flash addresses that pressure with a 552B-parameter multimodal Mixture-of-Experts model designed for contexts up to 1 million tokens. Its central wager is that aggressive cache compression and sparse retrieval can coexist with competitive agent performance.

Core Contribution

The paper combines a Causal Encoder-Decoder (CED) layout with Compressed Sparse Attention 2 (CSA2), FP4 KV caching, and deployment-oriented replay. During decoding, the model activates 16B parameters per token; during prefill it activates 8B. That asymmetry targets agent traces, where long inputs can dominate cost before a model produces much output.

The more consequential systems result is memory. The authors report a global KV-cache footprint of 890 bytes per token, about one quarter of DeepSeek-V4-Flash, and a persistent cache footprint roughly one eighth as large through SWA Bounded Replay. The distinction matters: global cache is retained in HBM, whereas persistent cache can reside on SSD or host memory. Reducing both addresses capacity and transfer bandwidth rather than shifting a single bottleneck elsewhere.

Technical Approach

The 40-layer network divides into a 20-layer causal encoder and a 20-layer decoder. Feed-forward blocks use DeepSeekMoE throughout. The first two encoder layers use sliding-window attention; later encoder layers use CSA2. In CSA2, a layer attends to a sparse set of context positions selected through an indexer rather than materializing and consuming full attention KV state at every layer.

Figure 3 lays out that encoder-decoder split and the placement of sliding-window attention, CSA2, Engram, DSpark, and the Hierarchical Sparse Indexer. The architecture is not simply a lower-precision cache: it reduces which KV states must be newly created, which can be reused across layers, and which positions are examined.

Figure 3. Overall architecture of DeepSeek-V4.1-Flash. The 40-layer network is divided into a causal encoder and a decoder, each with 20 layers. All feed-forward layers use standard DeepSeekMoE. The first two encoder layers use sliding window attention (SWA); the rest use Compressed Sparse Attention 2 (CSA2), with CSA2(ratio, mode) specifying the compression ratio and mode. The model also uses Single-Pass , Engram, DSpark, and a Hierarchical Sparse Indexer.

CSA2 has Full, Reindex, and Minimal modes. Full mode computes the main KV representation, indexer keys, and Top-K indices; Reindex recomputes selection while reusing earlier main KV and indexer keys; Minimal mode reuses all three except the current layer's query and sliding-window KV. The hierarchical indexer further constrains later selection: the decoder's first CSA2 Full layer chooses Top-512 indices and forms a shared candidate pool from selected blocks, while Reindex layers choose their Top-512 from that pool. This is a concrete way to make cache reuse compatible with changing queries across depth.

Results and Analysis

The authors pretrain on a 45T-token multimodal corpus and report post-training evaluations across text and multimodal agent settings. Figure 1 compares agentic benchmarks against named counterparts and plots the per-token global-cache reduction across DeepSeek generations. The reported approximately 4-fold reduction versus DeepSeek-V4-Flash and 437-fold reduction versus DeepSeek-V1 are large enough to matter operationally, particularly when serving many long-running sessions.

Figure 1. (a) Performance of DeepSeek-V4.1-Flash and its counterparts on agentic benchmarks. (b) Global KV cache size per token (in bytes) across generations of DeepSeek models, highlighting DeepSeek’s sustained efforts to reduce context memory requirements. DeepSeek-V4.1-Flash achieves approximately 4-fold and 437-fold reductions in per-token global KV cache size relative to DeepSeek-V4-Flash and DeepSeek-V1, respectively.

Figure 2 adds an important compute qualification. Its precision-weighted decode-FLOP plot shows nearly constant single-token decode cost as context length grows, using weights of 1 for BF16, 0.5 for FP8, and 0.25 for FP4 operations. That supports the paper's argument that the model attacks long-context inference at both the cache and decode-compute levels. The evidence is strongest for the authors' internal model comparison: the paper reports lower bits-per-byte on its held-out sets than both DeepSeek-V4-Flash-Base and DeepSeek-V4-Pro-Base, while agent curves improve as reinforcement-learning training scales and as maximal context reaches 1 million tokens.

The result is persuasive as a systems-and-model integration exercise, but the public evidence leaves several practical questions open. The headline memory ratios are relative to earlier DeepSeek models, not a cross-vendor serving study. The plots establish trends, yet the supplied material does not provide a cost-per-request analysis under production concurrency, cache misses, SSD traffic, or different hardware. The design should therefore be read as a substantial reduction in the authors' deployment footprint, not as a universal inference-cost ranking.

Evidence Box

moderate

Key Claims

  • Cross-layer CSA2 reuse and FP4 caching reduce long-context KV storage
  • CED lowers active parameters during prefill relative to decode
  • Sparse retrieval preserves agentic and multimodal capability under cache compression

Key Results

  • 890 bytes per token global KV cache, about 4× lower than DeepSeek-V4-Flash
  • Approximately 437× lower per-token global KV cache than DeepSeek-V1
  • 16B active parameters per decode token versus 8B during prefill
  • 1M-token maximum context and 45T-token multimodal pretraining corpus

Limitations & Caveats

  • Comparisons emphasize prior DeepSeek models rather than independent cross-vendor baselines
  • No supplied cost-per-request or production-concurrency evaluation
  • Persistent-cache savings depend on SWA Bounded Replay and SSD or host-memory deployment assumptions
  • Agent benchmark trends do not isolate the contribution of each compression component

Artifacts

Related Articles

Readers are encouraged to consult the original arXiv paper for complete details. SOTA Papers does not make claims beyond what is supported by the authors' reported evidence.