Fused KV Caching
-
Last Updated 9 August, 2026
-
by David Spuler, Ph.D.
What is Fused KV Caching?
Fused KV caching is merging two KV caches on the lengthwise token dimension. This means that the cache of two pieces of adjacent text can be created by just concatenating the two caches end-to-end. Surprisingly, this actually just works, but it did have some accuracy issues that needed research.
The nomenclature in this subarea of LLM optimization research is not yet settled, and the papers have used various different names for this technique:
- Fused KV caching
- Substring KV caching
- Concatenated KV caching
- Position-independent caching (PIC)
This has been an emerging area of research as a generalization of "prefix KV caching". The idea is to address the situation where the two pieces of text are not a prefix or a combined prefix. How do we combine two KV caches into one?
One particular situation where this commonly occurs is RAG chunks. To create the KV cache for two RAG chunks, we'd like to just pre-compute the KV cache for all the chunks, so we can just combine them in whatever order our reranker puts them. That would allow precomputed caches for the entire set of RAG chunks stored in the datastore alongside their text, and this would be available for all user queries. That speedup is the promise of this research.
This area of research, whatever you want to call it, just got a big boost from Meta research labs; see the "REFRAG" project by Lin et al. (2025). They didn't quite do the merging of two KV caches in the same way, but instead modified the attention algorithm to treat the chunks of text differently. The end result is very similar to precomputing and concatenating KV caches.
Fused KV Caching: Book Excerpts and Blog Articles
Free online book excerpts with full text chapters online and free PDF downloads, and the Aussie AI blog, including related articles:
- LLM prefix caching is an LLM inference optimization that is widely used in modern architectures. Any type of "cached tokens" is probably based on prefix KV caching. The idea is to store the KV cache for any prefix tokens in an LLM query, which occurs surprisingly commonly, such as ... more about LLM prefix caching »
- Prefix sharing is an LLM inference optimization that shares a common prefix across multiple queries. This is a generalization of KV prefix caching to the case where concurrent queries in a batch API have the same prefix. A common prefix can arise in a multi-query batch in common ways... more about Prefix sharing »
- KV caching is storing the results of the K and V vector operations that are performed in Transformer attention heads for LLM inference optimization. The data from the KV cache can be used to optimize the attention of the current query or multiple future queries ... more about KV caching »
- David Spuler, , September 26, 2024, RAG Optimization via Caching, https://www.aussieai.com/blog/rag-optimization-caching
- David Spuler, October 24, 2024, Generalizing Prefix KV Caching to RAG Chunks, Aussie AI Blog, https://www.aussieai.com/blog/prefix-kv-rag
- David Spuler, Ph.D., September 29, 2025, Promising LLM Inference Optimization Research, Aussie AI Blog, https://www.aussieai.com/blog/promising-llm-inference-optimization
- David Spuler, Michael Sharpe, June 2025, RAG Caching, Chapter 11, "RAG Optimization: Accurate and Efficient LLM Applications", https://www.aussieai.com/book/rag-book-11-rag-caching
- David Spuler, Ph.D., March 3rd, 2025, What's Hot in LLM Inference Optimization in 2025? Aussie AI Blog, https://www.aussieai.com/blog/hot-inference-optimization-2025
Research on Fused KV Caching
Research papers include:
- Jiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray, Yihua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu, Junchen Jiang, 3 Jun 2024 (v2), CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion, https://arxiv.org/abs/2405.16444 Code: https://github.com/YaoJiayi/CacheBlend.git (Generalizes prefix KV caching to KV cache fusion with selective recomputation of some KV cache data.)
- In Gim, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khandelwal, Lin Zhong, Nov 2023, Prompt Cache: Modular Attention Reuse for Low-Latency Inference, https://arxiv.org/abs/2311.04934 (Unique and insightful advance of generalizing KV caching to multiple prompts by computing a cache for short "segments" of prompts, including methods to adjust the different KV cache values for text segments that appear in different positions of the overall prompt.)
- Chao Jin, Zili Zhang, Xuanlin Jiang, Fangyue Liu, Xin Liu, Xuanzhe Liu, Xin Jin, 18 Apr 2024, RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation, https://arxiv.org/abs/2404.12457 (This paper briefly considers merging KV caches of multiple RAG chunks, but instead focuses on (a) caching of two or more chunks in one KV cache record, and (b) reordering the chunks in a cache-aware manner.)
- Hao Yu, Zelan Yang, Shen Li, Yong Li, Jianxin Wu, 11 Jun 2024, Effectively Compress KV Heads for LLM, https://arxiv.org/abs/2406.07056 (Examines KV cache head merging approaches for KV cache size reduction, and also examines RoPE encoding issues with relevance to fusing KV caches.)
- Zhongwei Wan, Ziang Wu, Che Liu, Jinfa Huang, Zhihong Zhu, Peng Jin, Longyue Wang, Li Yuan, 26 Jun 2024, LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference, https://arxiv.org/abs/2406.18139 (KV cache compression in text and multimodal inference, prioritizing eviction of text over image tokens, and using new ways to merge evicted KV cache data into the retained KV cache, including averaging, pivotal tokens, and weighted averages, which is relevant to token merging and KV cache fusion
- Yao Yao, Zuchao Li, Hai Zhao, 21 May 2024, SirLLM: Streaming Infinite Retentive LLM, https://arxiv.org/abs/2405.12528 (Low-rank decomposition to compress KV cache heads.)
- David Spuler, , September 26, 2024, RAG Optimization via Caching, https://www.aussieai.com/blog/rag-optimization-caching
- Songshuo Lu, Hua Wang, Yutian Rong, Zhi Chen, Yaohua Tang, 10 Oct 2024, TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text, https://arxiv.org/abs/2410.07590 (Fusing precomputed KV caches for each RAG chunk.)
- Junhao Hu, Wenrui Huang, Haoyi Wang, Weidong Wang, Tiancheng Hu, Qin Zhang, Hao Feng, Xusheng Chen, Yizhou Shan, Tao Xie, 20 Oct 2024, EPIC: Efficient Position-Independent Context Caching for Serving Large Language Models, https://arxiv.org/abs/2410.15332
- David Spuler, October 24, 2024, Generalizing Prefix KV Caching to RAG Chunks, Aussie AI Blog, https://www.aussieai.com/blog/prefix-kv-rag
- Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Monishwaran Maheswaran, June Paik, Michael W. Mahoney, Kurt Keutzer, Amir Gholami, 14 Nov 2024, Squeezed Attention: Accelerating Long Context Length LLM Inference, https://arxiv.org/abs/2411.09688 https://github.com/SqueezeAILab/SqueezedAttention (This is like a combination of semantic caching and prefix KV caching, and close to fused KV caching.)
- East Sun, Yan Wang, Lan Tian, 17 Oct 2024 (v4), Block-Attention for Efficient RAG, https://arxiv.org/abs/2409.15355
- Zhisong Zhang, Yan Wang, Xinting Huang, Tianqing Fang, Hongming Zhang, Chenlong Deng, Shuaiyi Li, Dong Yu, 21 Dec 2024, Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language Models, https://arxiv.org/abs/2412.16545 (Parallel encoding of chunks of context is similar to fused KV caching.)
- Philhoon Oh, Jinwoo Shin, James Thorne, 13 Jan 2025, Parallel Key-Value Cache Fusion for Position Invariant RAG, https://arxiv.org/abs/2501.07523 (Generating the KV cache for each RAG chunk.)
- Longze Chen, Jan 2025 (accessed), Awesome-KV-Cache-Compression: Must-read papers on KV Cache Compression (constantly updating), https://github.com/October2001/Awesome-KV-Cache-Compression (KV cache reuse across multiple prompts via SwiftKV, sounds similar to prefix KV caching or fused KV caching, and also SingleInputKV does KV cache layer fusion in a single prompt.)
- Shiju Zhao, Junhao Hu, Rongxiao Huang, Jiaqi Zheng, Guihai Chen, 4 Feb 2025, MPIC: Position-Independent Multimodal Context Caching System for Efficient MLLM Serving, https://arxiv.org/abs/2502.01960
- Kun Luo, Zheng Liu, Peitian Zhang, Hongjin Qian, Jun Zhao, Kang Liu, 17 Feb 2025, Does RAG Really Perform Bad For Long-Context Processing? https://arxiv.org/abs/2502.11444 (Long context RAG processing based on the KV cache data is similar to fused/substring KV caching methods.)
- S Agarwal, S Sundaresan, S Mitra, D Mahapatra, Feb 2025, Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation, https://skejriwal44.github.io/docs/CacheCraft_SIGMOD_2025.pdf (Managing pre-computed KV caches for RAG chunks as a generalization of prefix KV caching, addressing limitations in their position and ordering.)
- Jingbo Yang, Bairu Hou, Wei Wei, Yujia Bao, Shiyu Chang, 21 Feb 2025, KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse, https://arxiv.org/abs/2502.16002 https://github.com/UCSB-NLP-Chang/KVLink (Computing a KV cache for each RAG chunk, and using techniques to fuse/merge/concatenate these KV caches, i.e., fused KV caching as a generalization of prefix KV caching, while restoring cross-chunk attention accuracy via 3 techniques: positional re-encoding, "link tokens" between chunks processed during inference, and fine-tuning).
- Shai Bergman, Zhang Ji, Anne-Marie Kermarrec, Diana Petrescu, Rafael Pires, Mathis Randl, Martijn de Vos, 7 Mar 2025, Leveraging Approximate Caching for Faster Retrieval-Augmented Generation, https://arxiv.org/abs/2503.05530
- J Hu, W Huang, W Wang, H Wang, H Feng, X Chen, 2025, EPIC: Efficient Position-Independent Caching for Serving Large Language Models, Proceedings of the 42nd International Conference on Machine Learning, Vancouver, Canada. PMLR 267, 2025, https://openreview.net/pdf?id=qjd3ZUiHRT
- Xiaoqiang Lin, Aritra Ghosh, Bryan Kian Hsiang Low, Anshumali Shrivastava, Vijai Mohan, 1 Sep 2025, REFRAG: Rethinking RAG based Decoding, https://www.arxiv.org/abs/2509.01092 https://www.alphaxiv.org/pdf/2509.01092 (Separates the attention computations across RAG chunks, which is effectively the same as "fused KV" or "concatenated KV" approaches with pre-computed per-chunk KV caches.)
- Jiahao Wang, Weiyu Xie, Mingxing Zhang, Boxing Zhang, Jianwei Dong, Yuening Zhu, Chen Lin, Jinqi Tang, Yaochen Han, Zhiyuan Ai, Xianglin Chen, Yongwei Wu, Congfeng Jiang, 19 Jan 2026, From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation, https://arxiv.org/abs/2601.12904 (Improve accuracy of concatenating fused KV caches by re-computing some of the caches for tokens that appear in the main query.)
- Shiju Zhao, Junhao Hu, Jiaqi Zheng, Guihai Chen, 2 Feb 2026, You Need an Encoder for Native Position-Independent Caching, https://arxiv.org/abs/2602.01519 https://github.com/shijuzhao/Comb
- Yang Liu and Yunfei Gu, Shanghai Jiao Tong University; Liqiang Zhang, Jinan Inspur Data Technology Co., Ltd; Chentao Wu, Guangtao Xue, Jie Li, and Minyi Guo, Shanghai Jiao Tong University; Junhao Hu, Peking University; Jie Meng, Huawei Cloud, Feb 2026, CacheSlide: Unlocking Cross Position-Aware KV Cache Reuse for Accelerating LLM Serving, Proceedings of the 24th USENIX Conference on File and Storage Technologies. February 24–26, 2026 • Santa Clara, CA, USA, https://www.usenix.org/system/files/fast26-liu-yang.pdf
- Samuel Cestola, Tianxiang Xia, Zheng Weiyan, Zheng Pengfei, Diego Didona, 3 Mar 2026, An experimental study of KV cache reuse strategies in chunk-level caching systems, https://arxiv.org/abs/2603.20218
- David Spuler, Ph.D., September 29, 2025, Promising LLM Inference Optimization Research, Aussie AI Blog, https://www.aussieai.com/blog/promising-llm-inference-optimization
- David Spuler, Michael Sharpe, June 2025, RAG Caching, Chapter 11, "RAG Optimization: Accurate and Efficient LLM Applications", https://www.aussieai.com/book/rag-book-11-rag-caching
- David Spuler, Ph.D., March 3rd, 2025, What's Hot in LLM Inference Optimization in 2025? Aussie AI Blog, https://www.aussieai.com/blog/hot-inference-optimization-2025
- Chuangtao Chen, Grace Li Zhang, Xunzhao Yin, Cheng Zhuo, Bing Li, Ulf Schlichtmann, 17 Apr 2026 (v2), KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs, https://arxiv.org/abs/2604.13226
- Yingsheng Geng, Yuchong Gao, Weihong Wu, Guyue Liu, Jiang Liu, 28 Feb 2026, RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache Reuse, https://arxiv.org/abs/2603.13289
- Zhenyu He, Jun Zhang, Shengjie Luo, Jingjing Xu, Zhi Zhang, Di He, 4 Mar 2025 (v2), Let the Code LLM Edit Itself When You Edit the Code, https://arxiv.org/abs/2407.03157 https://github.com/zhenyuhe00/PIE (Correcting KV caching by removing RoPE and re-applying it.)
- Junlong Tong, Jinlan Fu, Zixuan Lin, Yingqi Fan, Anhao Zhao, Hui Su, and Xiaoyu Shen. 2025. LLM as Effective Streaming Processor: Bridging Streaming-Batch Mismatches with Group Position Encoding. In Findings of the Association for Computational Linguistics: ACL 2025, pages 23497–23517, Vienna, Austria. Association for Computational Linguistics. https://aclanthology.org/2025.findings-acl.1207/ https://aclanthology.org/2025.findings-acl.1207.pdf
- Qingyao Li, Wei Xia, Xinyi Dai, Kounianhua Du, Weiwen Liu, Yasheng Wang, Ruiming Tang, Yong Yu, and Weinan Zhang. 2025. RethinkMCTS: Refining Erroneous Thoughts in Monte Carlo Tree Search for Code Generation. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 8092–8110, Suzhou, China. Association for Computational Linguistics. https://aclanthology.org/2025.emnlp-main.410/ https://aclanthology.org/2025.emnlp-main.410.pdf
- Yuechi Zhou, Yi Su, Jianxin Zhang, Juntao Li, Qingrong Xia, Zhefeng Wang, Xinyu Duan, Baoxing Huai, 13 Nov 2025, A3: Attention-Aware Accurate KV Cache Fusion for Fast Large Language Model Serving, https://arxiv.org/abs/2511.17560
- Xize Cheng, Dongjie Fu, Chenyuhao Wen, Tao Jin, Hai Yu, Di Cao, Qinying Liu, Yexin Yang, Zehan Wang, Shengpeng Ji, Siqi Zheng, Xu Tan, Zhou Zhao, 11 Feb 2026 (updated), Vox-Infinity: Benchmarking the Limits of Long-Context Spoken Language Models, Submitted to ICLR 2026, https://openreview.net/forum?id=6dKwqnT7bu https://openreview.net/pdf?id=6dKwqnT7bu https://vox-infinity.github.io/
- Xin Teng, Canyu Zhang, Shaoyi Zheng, Danyang Zhuo, Tianyi Zhou, Shengjie Wang, 5 Mar 2026, InfoFlow KV: Information-Flow-Aware KV Recomputation for Long Context, https://arxiv.org/abs/2603.05353 (Merging KV caches of chunks via selective per-token recomputation.)
- Shihao Wang, Jiahao Chen, Yanqi Pan, Hao Huang, Yichen Hao, Xiangyu Zou, Wen Xia, Wentao Zhang, Chongyang Qiu, Pengfei Wang, 5 Feb 2026 (v3), ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation, https://arxiv.org/abs/2602.02579
- Hancheng Ye, Zhengqi Gao, Mingyuan Ma, Qinsi Wang, Yuzhe Fu, Ming-Yu Chung, Yueqian Lin, Zhijian Liu, Jianyi Zhang, Danyang Zhuo, Yiran Chen, 1 Nov 2025 (v2), KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems, https://arxiv.org/abs/2510.12872
- Jie Ou, Jinyu Guo, Shiyao Guo, Yuang Li, Ruiqi Wu, Zhaokun Wang, Wenyi Li, Wenhong Tian, 5 May 2026, AdapShot: Adaptive Many-Shot In-Context Learning with Semantic-Aware KV Cache Reuse, https://arxiv.org/abs/2605.03644
- Rya Sanovar, Srikant Bharadwaj, Hritvik Taneja, Moinuddin Qureshi, 8 Jun 2026, SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance, https://arxiv.org/abs/2606.09441
- Jiahao Wang, Weiyu Xie, Mingxing Zhang, Boxin Zhang, Jianwei Dong, Yuening Zhu, Chen Lin, Jingqi Tang, Yaochen Han, Zhiyuan Ai, Xianglin Chen, Yongwei Wu, and Congfeng Jiang. 2026. From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation. Proc. ACM Manag. Data 4, 1, Article 41 (February 2026), 28 pages. https://doi.org/10.1145/3786655 https://dl.acm.org/doi/abs/10.1145/3786655
- Yinsicheng Jiang, Yeqi Huang, Liang Cheng, Cheng Deng, Xuan Sun, Luo Mai, 6 May 2026 (v4), ContextPilot: Fast Long-Context Inference via Context Reuse, https://arxiv.org/abs/2511.03475
- Haocheng Xia, Mihir Pamnani, Hanxi Fang, Supawit Chockchowwat, Yongjoo Park, 3 Jun 2026, LazyAttention: Efficient Retrieval-Augmented Generation with Deferred Positional Encoding, https://arxiv.org/abs/2606.04302
- Jinbo Su, Yuxuan Hu, Cuiping Li, Hong Chen, Jia Li, Lintao Ma, Jing Zhang, 13 Jan 2026, TableCache: Primary Foreign Key Guided KV Cache Precomputation for Low Latency Text-to-SQL, https://arxiv.org/abs/2601.08743
More AI Research Topics
Read more about:
- 500+ LLM Inference Optimization Techniques
- What's Hot in LLM Inference Optimization in 2025?
- Inference Optimization Research
- « Research Home