Skip to content
AI Atlas
Papercs.CL

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

Published 18 Sept 2026arXiv:2609.19969

data quality78

Updated 4 h ago · first seen 17 Sept 2026

paper_01M2SE0KQKZSK96683D0DN8FS8

Abstract

The widespread adoption of long-horizon agents has made model workloads increasingly input-heavy. Although prior work has substantially reduced the cost of long-context computation, prefill remains computationally expensive, and large KV caches continue to strain HBM and SSD capacity and data-transfer bandwidth. Together, these compute, storage, and bandwidth demands constitute the primary bottleneck to further lowering deployment costs. To address this challenge, we introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. With its Causal Encoder-Decoder (CED) architecture, the model activates 16B parameters per token during decode but only 8B parameters during prefill, substantially improving cost efficiency for agentic workloads. To push the limits of KV cache compression, DeepSeek-V4.1-Flash combines cross-layer KV cache reuse in Compressed Sparse Attention 2 (CSA2) with FP4 KV caching. These designs reduce its global KV cache footprint (always in HBM) to 890 bytes per token, roughly 1/4 of the corresponding footprint of DeepSeek-V4-Flash. Further, through a dedicated deployment optimization known as SWA Bounded Replay, DeepSeek-V4.1-Flash reduces its persistent KV cache footprint (always on SSD or in host memory) to roughly 1/8 of that of DeepSeek-V4-Flash. Despite its much smaller KV cache footprint, the model delivers substantially better performance than the baseline. In addition, we streamline the DeepSeek-V4 architecture and introduce several efficient architectural extensions. We pretrain DeepSeek-V4.1-Flash on a multimodal corpus comprising 45T tokens and conduct comprehensive post-training, yielding strong performance across diverse text-based and multimodal agentic scenarios. Model checkpoints are available at https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash.

Authors

Authors 100

:Anyi XuB. LiBangcai LinBing XueBingCheng XianBingzheng XuBochao WuBowei ZhangBoyi DengC. C. YuChao JinChaofan LinChen DongChenbing WangChenfan FengChengda LuChenggang ZhaoChengqi DengChengyuan ZhangChenhao XuChenqi ZhaoChenze ShaoChuhao WangChuqi ZhangDamai DaiDeepSeek-AIDejian YangDeli ChenDi HuangDi WuDonghao LiErhang LiEric FuF. ZhouFangwei ZhouFangyun LinFangzhou YuanFeiyu XiaFucong DaiGuangbo HaoGuanglin LiGuanting ChenGuoai CaoGuofan FanGuolai MengGuowei LiHaichuan ZhangHaiyang MaHaiyang ShenHan LiHan YuHan ZhangHangyuan DengHanwei XuHanxiang XuHanxun ZhongHao GuoHao JiangHao LiHao QinHaodong WenHaofen LiangHaofeng HuangHaohua LiuHaoling ZhangHaoming LuoHaoran YangHaotian XuHaotian YuanHaoting HuangHaowen LuoHaoyang CaiHaoyu ChenHaozhe JiHengran ZhangHengrui WangHengxu WuHonghui DingHongxuan TangHuadong WangHuanqi CaoHuazuo GaoHui QuHui ZengJ. H. JinJ. H. ZhangJ. X. ZouJ. YangJia YuJiahui ZhouJiajun ChenJialiang HuangJialin ZhaoJiamin TangJian ZhouJianan TongJianwen LiJiaqi ZhuJiarui Wang

Linked names open researcher pages (created from the paper's author list; name-only, no affiliation unless a source states it). Unlinked names have no researcher record yet.

Organizations

Organizations 0

No organization stated. arXiv metadata does not carry affiliations; an organization is linked only when a model card or lab page cites the paper.

Models

Models introduced or described 0

Inbound described_by relations from model cards and documentation.

No model links this paper yet

Model pages link papers through their model cards and documentation; the relation is written only when a source states it.

Datasets

Datasets used 0

No dataset relation recorded.

Benchmarks

Benchmarks used 0

No benchmark relation recorded.

Code

Repositories & frameworks 0

No repository linked.

Timeline

Timeline 3

  • DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression: published at changed from 2026-09-17T00:00:00+00:00 to 2026-09-18T04:00:00+00:00

    Published17 Sept 202618 Sept 2026arxiv
  • DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression: authors changed from ["Anyi Xu", "B. Li", "Bangcai Lin", "Bing Xue", "BingChen… to [":", "Anyi Xu", "B. Li", "Bangcai Lin", "Bing Xue", "Bin…

    AuthorsAnyi Xu, B. Li, Bangcai Lin, Bing Xue, BingCheng Xian, Bingzheng Xu, Bochao Wu, Bowei Zhang, Boyi Deng, C. C. Yu, Chao Jin, Chaofan Lin, Chen Dong, Chenbing Wang, Chenfan Feng, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chengyuan Zhang, Chenhao Xu, Chenqi Zhao, Chenze Shao, Chuhao Wang, Chuqi Zhang, Damai Dai, DeepSeek-AI, Dejian Yang, Deli Chen, Di Huang, Di Wu, Donghao Li, Erhang Li, Eric Fu, F. Zhou, Fangwei Zhou, Fangyun Lin, Fangzhou Yuan, Feiyu Xia, Fucong Dai, Guangbo Hao, Guanglin Li, Guanting Chen, Guoai Cao, Guofan Fan, Guolai Meng, Guowei Li, Haichuan Zhang, Haiyang Ma, Haiyang Shen, Han Li:, Anyi Xu, B. Li, Bangcai Lin, Bing Xue, BingCheng Xian, Bingzheng Xu, Bochao Wu, Bowei Zhang, Boyi Deng, C. C. Yu, Chao Jin, Chaofan Lin, Chen Dong, Chenbing Wang, Chenfan Feng, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chengyuan Zhang, Chenhao Xu, Chenqi Zhao, Chenze Shao, Chuhao Wang, Chuqi Zhang, Damai Dai, DeepSeek-AI, Dejian Yang, Deli Chen, Di Huang, Di Wu, Donghao Li, Erhang Li, Eric Fu, F. Zhou, Fangwei Zhou, Fangyun Lin, Fangzhou Yuan, Feiyu Xia, Fucong Dai, Guangbo Hao, Guanglin Li, Guanting Chen, Guoai Cao, Guofan Fan, Guolai Meng, Guowei Li, Haichuan Zhang, Haiyang Ma, Haiyang Shen, Han Li, Han Yu, Han Zhang, Hangyuan Deng, Hanwei Xu, Hanxiang Xu, Hanxun Zhong, Hao Guo, Hao Jiang, Hao Li, Hao Qin, Haodong Wen, Haofen Liang, Haofeng Huang, Haohua Liu, Haoling Zhang, Haoming Luo, Haoran Yang, Haotian Xu, Haotian Yuan, Haoting Huang, Haowen Luo, Haoyang Cai, Haoyu Chen, Haozhe Ji, Hengran Zhang, Hengrui Wang, Hengxu Wu, Honghui Ding, Hongxuan Tang, Huadong Wang, Huanqi Cao, Huazuo Gao, Hui Qu, Hui Zeng, J. H. Jin, J. H. Zhang, J. X. Zou, J. Yang, Jia Yu, Jiahui Zhou, Jiajun Chen, Jialiang Huang, Jialin Zhao, Jiamin Tang, Jian Zhou, Jianan Tong, Jianwen Li, Jiaqi Zhu, Jiarui Wangarxiv
  • New paper: DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

    huggingface

Full timeline →

Sources

Sources 2

Source documents
SourceDocumentTypeTierLast observedSnapshots
arXiv (Atom API + RSS)rss.arxiv.org/rss/cs.CL feedT1· Official4 h ago7
Hugging Face Hub (public pages, model cards, papers)huggingface.co/papers listingT2· Quality secondary5 h ago40

Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.