KV Cache Correction

  • Last Updated 9 August, 2026
  • by David Spuler, Ph.D.

KV Cache Correction: Book Excerpts and Blog Articles

Free online book excerpts with full text chapters online and free PDF downloads, and the Aussie AI blog, including related articles:

  • Partial RoPE is an optimization of Rotational Positional Encoding where only a subset of the vectors are rotated. RoPE is a type of relative positional encoding inside the attention kernel. The idea of partial RoPE is that rotation primarily helps focus on nearby tokens, and the meaning of tokens that are far away is diminished. To increase the effect of distant but important tokens or facts in long contexts, some tokens are left unrotated. This method is primarily used for improvement of the performance and accuracy of the attention module in long contexts, but it also gives a minor improvement in ... more about Partial RoPE »
  • David Spuler, May 31st, 2026, Chapter 32. Positional Encoding, in book LLM Inference Optimization: State-of-the-Art Research, Table of Contents: https://www.aussieai.com/book/llm-inference-optimization https://www.amazon.com/dp/B0H3FKR39T

Research on KV Cache Correction

Research papers include:

  • Chuangtao Chen, Grace Li Zhang, Xunzhao Yin, Cheng Zhuo, Bing Li, Ulf Schlichtmann, 17 Apr 2026 (v2), KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs, https://arxiv.org/abs/2604.13226
  • Yingsheng Geng, Yuchong Gao, Weihong Wu, Guyue Liu, Jiang Liu, 28 Feb 2026, RelayCaching: Accelerating LLM Collaboration via Decoding KV Cache Reuse, https://arxiv.org/abs/2603.13289
  • Zhenyu He, Jun Zhang, Shengjie Luo, Jingjing Xu, Zhi Zhang, Di He, 4 Mar 2025 (v2), Let the Code LLM Edit Itself When You Edit the Code, https://arxiv.org/abs/2407.03157 https://github.com/zhenyuhe00/PIE (Correcting KV caching by removing RoPE and re-applying it.)
  • Ye Qiao, Haocheng Xu, Xiaofan Zhang, Sitao Huang, 26 Sep 2025, Rethinking RoPE Scaling in Quantized LLM: Theory, Outlier, and Channel-Band Analysis with Weight Rescaling, https://arxiv.org/abs/2510.00028
  • Junu Kim, Xiao Liu, Zhenghao Lin, Lei Ji, Yeyun Gong, Edward Choi, 14 Nov 2025 (updated), Behind RoPE: How Does Causal Mask Encode Positional Information? ICLR 2026 Conference Withdrawn Submission, https://openreview.net/forum?id=IAXBLI2vo5 https://openreview.net/pdf?id=IAXBLI2vo5
  • Qiao, Y., & Huang, S. (2026). Q-ROAR: Outlier-Aware Rescaling for RoPE Position Interpolation in Quantized Long-Context LLMs (Student Abstract). Proceedings of the AAAI Conference on Artificial Intelligence, 40(48), 41359-41361. https://doi.org/10.1609/aaai.v40i48.42269 https://ojs.aaai.org/index.php/AAAI/article/view/42269
  • Xin Teng, Canyu Zhang, Shaoyi Zheng, Danyang Zhuo, Tianyi Zhou, Shengjie Wang, 5 Mar 2026, InfoFlow KV: Information-Flow-Aware KV Recomputation for Long Context, https://arxiv.org/abs/2603.05353 (Merging KV caches of chunks via selective per-token recomputation.)
  • Shihao Wang, Jiahao Chen, Yanqi Pan, Hao Huang, Yichen Hao, Xiangyu Zou, Wen Xia, Wentao Zhang, Chongyang Qiu, Pengfei Wang, 5 Feb 2026 (v3), ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation, https://arxiv.org/abs/2602.02579
  • Aurick Qiao, Zhewei Yao, Samyam Rajbhandari, and Yuxiong He. 2025. SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 25734–25753, Suzhou, China. Association for Computational Linguistics. https://aclanthology.org/2025.emnlp-main.1306/ https://github.com/snowflakedb/arctictraining and https://github.com/snowflakedb/arcticinference (Prefill-only KV layer fusion with propagation of KV caches to later layers allows prefill layer skipping.)
  • Shoaib Ahmed Siddiqui, Xin Dong, Greg Heinrich, Thomas Breuel, Jan Kautz, David Krueger, Pavlo Molchanov, 23 Jul 2024, A deeper look at depth pruning of LLMs, https://arxiv.org/abs/2407.16286
  • Haoyi Wu, Kewei Tu, 17 May 2024, Layer-Condensed KV Cache for Efficient Inference of Large Language Models, https://arxiv.org/abs/2405.10637 Code: https://github.com/whyNLP/LCKV (Use the KV cache for only the final layer as the KV cache for all other layers, or alternatively, use only the cache from a few layers, also possibly using a few standard layers as "warmup layers". This idea is conceptuatlly similar to "propagation" of the KV cache in early exit methods or to layer fusion of weights.)
  • Ruijie Miao, Yihan Yan, Xinshuo Yao, Tong Yang, 25 Jul 2024, An Efficient Inference Framework for Early-exit Large Language Models, https://arxiv.org/abs/2407.20272 (Faster early exit using batching and KV cache resolution.)
  • Sangmin Yoo, Srikanth Malla, Chiho Choi, Wei D. Lu, Joon Hee Choi, 7 Jan 2026, ADEPT: Adaptive Dynamic Early-Exit Process for Transformers, https://arxiv.org/abs/2601.03700 (Early exit of layers during prefill with a mechanism to fix the KV cache for missing layers when later needed in the decoding phase.)
  • Meng, L., Zhang, R., Shan, W. (2026). Robust and Efficient Early Exit for Large Language Models: Mitigating KV Cache Loss and Enhancing Exit Stability. In: Jin, L., Wang, L. (eds) Advances in Neural Networks – ISNN 2025. ISNN 2025. Lecture Notes in Computer Science, vol 15951. Springer, Singapore. https://doi.org/10.1007/978-981-95-1233-1_7 https://link.springer.com/chapter/10.1007/978-981-95-1233-1_7
  • Yuhan Liu, Yuyang Huang, Jiayi Yao, Shaoting Feng, Zhuohan Gu, Kuntai Du, Hanchen Li, Yihua Cheng, and Junchen Jiang, Shan Lu, Madan Musuvathi, and Esha Choukse, May 2026, DroidSpeak: KV Cache Sharing Across Fine-tuned Model Variants, Proceedings of the 23rd USENIX Symposium on Networked Systems Design and Implementation. May 4–6, 2026 • Renton, WA, USA, https://www.usenix.org/system/files/nsdi26-liu-yuhan.pdf https://www.usenix.org/conference/nsdi26/presentation/liu-yuhan (Sharing prefix KV caches from different fine-tuned LLMs, with partial layerwise recomputation.)
  • David Spuler, May 31st, 2026, Chapter 32. Positional Encoding, in book LLM Inference Optimization: State-of-the-Art Research, Table of Contents: https://www.aussieai.com/book/llm-inference-optimization https://www.amazon.com/dp/B0H3FKR39T
  • Utkarsh Saxena, Kaushik Roy, 6 Oct 2025, KVLinC : KV Cache Quantization with Hadamard Rotation and Linear Correction, https://arxiv.org/abs/2510.05373
  • Manan Gupta, Dhruv Kumar, 20 Apr 2026, Latent Phase-Shift Rollback: Inference-Time Error Correction via Residual Stream Monitoring and KV-Cache Steering, https://arxiv.org/abs/2604.18567
  • Haocheng Xia, Mihir Pamnani, Hanxi Fang, Supawit Chockchowwat, Yongjoo Park, 3 Jun 2026, LazyAttention: Efficient Retrieval-Augmented Generation with Deferred Positional Encoding, https://arxiv.org/abs/2606.04302

More AI Research Topics

Read more about: