KV Cache Reversal
-
Last Updated 9 August, 2026
-
by David Spuler, Ph.D.
KV Cache Reversal: Book Excerpts and Blog Articles
Free online book excerpts with full text chapters online and free PDF downloads, and the Aussie AI blog, including related articles:
- Partial RoPE is an optimization of Rotational Positional Encoding where only a subset of the vectors are rotated. RoPE is a type of relative positional encoding inside the attention kernel. The idea of partial RoPE is that rotation primarily helps focus on nearby tokens, and the meaning of tokens that are far away is diminished. To increase the effect of distant but important tokens or facts in long contexts, some tokens are left unrotated. This method is primarily used for improvement of the performance and accuracy of the attention module in long contexts, but it also gives a minor improvement in ... more about Partial RoPE »
- David Spuler, May 31st, 2026, Chapter 32. Positional Encoding, in book LLM Inference Optimization: State-of-the-Art Research, Table of Contents: https://www.aussieai.com/book/llm-inference-optimization https://www.amazon.com/dp/B0H3FKR39T
Research on KV Cache Reversal
Research papers include:
- Zhenyu He, Jun Zhang, Shengjie Luo, Jingjing Xu, Zhi Zhang, Di He, 4 Mar 2025 (v2), Let the Code LLM Edit Itself When You Edit the Code, https://arxiv.org/abs/2407.03157 https://github.com/zhenyuhe00/PIE (Correcting KV caching by removing RoPE and re-applying it.)
- Ye Qiao, Haocheng Xu, Xiaofan Zhang, Sitao Huang, 26 Sep 2025, Rethinking RoPE Scaling in Quantized LLM: Theory, Outlier, and Channel-Band Analysis with Weight Rescaling, https://arxiv.org/abs/2510.00028
- Junu Kim, Xiao Liu, Zhenghao Lin, Lei Ji, Yeyun Gong, Edward Choi, 14 Nov 2025 (updated), Behind RoPE: How Does Causal Mask Encode Positional Information? ICLR 2026 Conference Withdrawn Submission, https://openreview.net/forum?id=IAXBLI2vo5 https://openreview.net/pdf?id=IAXBLI2vo5
- Qiao, Y., & Huang, S. (2026). Q-ROAR: Outlier-Aware Rescaling for RoPE Position Interpolation in Quantized Long-Context LLMs (Student Abstract). Proceedings of the AAAI Conference on Artificial Intelligence, 40(48), 41359-41361. https://doi.org/10.1609/aaai.v40i48.42269 https://ojs.aaai.org/index.php/AAAI/article/view/42269
- Xin Teng, Canyu Zhang, Shaoyi Zheng, Danyang Zhuo, Tianyi Zhou, Shengjie Wang, 5 Mar 2026, InfoFlow KV: Information-Flow-Aware KV Recomputation for Long Context, https://arxiv.org/abs/2603.05353 (Merging KV caches of chunks via selective per-token recomputation.)
- David Spuler, May 31st, 2026, Chapter 32. Positional Encoding, in book LLM Inference Optimization: State-of-the-Art Research, Table of Contents: https://www.aussieai.com/book/llm-inference-optimization https://www.amazon.com/dp/B0H3FKR39T
- Haocheng Xia, Mihir Pamnani, Hanxi Fang, Supawit Chockchowwat, Yongjoo Park, 3 Jun 2026, LazyAttention: Efficient Retrieval-Augmented Generation with Deferred Positional Encoding, https://arxiv.org/abs/2606.04302
More AI Research Topics
Read more about:
- 500+ LLM Inference Optimization Techniques
- What's Hot in LLM Inference Optimization in 2025?
- Inference Optimization Research
- « Research Home