Aussie AI
Infinite Context Inference
-
Last Updated 19 June, 2026
-
by David Spuler, Ph.D.
Infinite Context Inference: Book Excerpts and Blog Articles
Free online book excerpts with full text chapters online and free PDF downloads, and the Aussie AI blog, including related articles:
- David Spuler, May 31st, 2026, Chapter 43. Long, Ultralong and Infinite Context, in book LLM Inference Optimization: State-of-the-Art Research, Table of Contents: https://www.aussieai.com/book/llm-inference-optimization https://www.amazon.com/dp/B0H3FKR39T
Research on Infinite Context Inference
Research papers include:
- Dan McAteer, May 9, 2026, Anthropic is hinting at “infinite” context windows, https://x.com/daniel_mac8/status/2052350232688832763
- Bin Lin, Chen Zhang, Tao Peng, Hanyu Zhao, Wencong Xiao, Minmin Sun, Anmin Liu, Zhipeng Zhang, Lanbo Li, Xiafei Qiu, Shen Li, Zhigang Ji, Tao Xie, Yong Li, Wei Lin, 4 Jul 2024 (v2), Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache, https://arxiv.org/abs/2401.02669
- Chien Van Nguyen, Huy Nguyen, Ruiyi Zhang, Hanieh Deilamsalehy, Puneet Mathur, Viet Dac Lai, Haoliang Wang, Jayakumar Subramanian, Ryan A. Rossi, Trung Bui, Nikos Vlassis, Franck Dernoncourt, Thien Huu Nguyen, 18 Apr 2026 (v4), Lizard: An Efficient Linearization Framework for Large Language Models, https://arxiv.org/abs/2507.09025
- Pai Zeng, Zhenyu Ning, Jieru Zhao, Weihao Cui, Mengwei Xu, Liwei Guo, Xusheng Chen, Yizhou Shan, 27 May 2024 (v2), The CAP Principle for LLM Serving: A Survey of Long-Context Large Language Model Serving, https://arxiv.org/abs/2405.11299
- Surendra Pathak, Bo Han, 11 Apr 2026 (v2), Towards Efficient Large Vision-Language Models: A Comprehensive Survey on Inference Strategies, https://arxiv.org/abs/2603.27960
- Bingyang Wu, Shengyu Liu, Yinmin Zhong, Peng Sun, Xuanzhe Liu, and Xin Jin. 2024. LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism. In Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles (SOSP '24). Association for Computing Machinery, New York, NY, USA, 640–654. https://doi.org/10.1145/3694715.3695948 https://dl.acm.org/doi/abs/10.1145/3694715.3695948
- Chaojun Xiao, Pengle Zhang, Xu Han, Guangxuan Xiao, Yankai Lin, Zhengyan Zhang, Zhiyuan Liu, Maosong Sun, 28 May 2024 (v2), InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory, https://arxiv.org/abs/2402.04617 https://github.com/thunlp/InfLLM
- Heejun Lee, Geon Park, Jaduk Suh, Sung Ju Hwang, 13 Feb 2025, InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU, https://arxiv.org/abs/2502.08910
- Aloy Banerjee, May 7, 2025, Infinite Context Length in LLMs — The Next Big Advantage in AI, https://medium.com/@aloy.banerjee30/infinite-context-length-in-llms-the-next-big-advantage-in-ai-2550e9e6ce9b
- Tsendsuren Munkhdalai, Manaal Faruqui, Siddharth Gopal, 9 Aug 2024 (v2), Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention, https://arxiv.org/abs/2404.07143
- Amirkeivan Mohtashami, Martin Jaggi, 20 Nov 2023 (v2), Landmark Attention: Random-Access Infinite Context Length for Transformers, https://arxiv.org/abs/2305.16300
- Hao Liu, Matei Zaharia, Pieter Abbeel, 27 Nov 2023 (v4), Ring Attention with Blockwise Transformers for Near-Infinite Context, https://arxiv.org/abs/2310.01889
- Won-Gi Paeng, Daesuk Kwon, Kyungwon Jeong, Honggyo Suh, 1 May 2025 (v5), Folded Context Condensation in Path Integral Formalism for Infinite Context Transformers, https://arxiv.org/abs/2405.04620
- Ziming Liu, Shaoyu Wang, Shenggan Cheng, Zhongkai Zhao, Kai Wang, Xuanlei Zhao, James Demmel, Yang You, 28 Sep 2025 (v4), StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training, https://arxiv.org/abs/2407.00611
- Zafeirios Fountas, Martin A Benfeghoul, Adnan Oomerjee, Fenia Christopoulou, Gerasimos Lampouras, Haitham Bou-Ammar, Jun Wang, 10 Oct 2025 (v3), Human-inspired Episodic Memory for Infinite Context LLMs, https://arxiv.org/abs/2407.09450
- Xiaoran Liu, Ruixiao Li, Qipeng Guo, Zhigeng Liu, Yuerong Song, Kai Lv, Hang Yan, Linlin Li, Qun Liu, Xipeng Qiu, 19 Mar 2025 (v3), ReAttention: Training-Free Infinite Context with Finite Attention Scope, https://arxiv.org/abs/2407.15176
- Minsoo Kim, Kyuhong Shim, Jungwook Choi, Simyung Chang, 2 Oct 2024, InfiniPot: Infinite Context Processing on Memory-Constrained LLMs, https://arxiv.org/abs/2410.01518
- Zongwu Wang, Fangxin Liu, Mingshuai Li, Li Jiang, 29 Dec 2024, TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication, https://arxiv.org/abs/2412.20501
- Yang Zhou, Hongyi Liu, Zhuoming Chen, Yuandong Tian, Beidi Chen, 7 Feb 2025, GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? https://arxiv.org/abs/2502.05252
- Xiaoju Ye, Zhichun Wang, Jingyuan Wang, 18 Feb 2025, Infinite Retrieval: Attention Enhanced LLMs in Long-Context Processing, https://arxiv.org/abs/2502.12962
- Jiyu Chen, Shuang Peng, Daxiong Luo, Fan Yang, Renshou Wu, Fangyuan Li, Xiaoxin Chen, 28 Mar 2025, EdgeInfinite: A Memory-Efficient Infinite-Context Transformer for Edge Devices, https://arxiv.org/abs/2503.22196
- Tao An, 8 Aug 2025, Cognitive Workspace: Active Memory Management for LLMs -- An Empirical Study of Functional Infinite Context, https://arxiv.org/abs/2508.13171
- Oliver Zahn, Matt Beton, Simran Chana, 4 Feb 2026 (v2), Attention Is Not Retention: The Orthogonality Constraint in Infinite-Context Architectures, https://arxiv.org/abs/2601.15313
- Yushi Bai, Qian Dong, Ting Jiang, Xin Lv, Zhengxiao Du, Aohan Zeng, Jie Tang, Juanzi Li, 12 Mar 2026, IndexCache: Accelerating Sparse Attention via Cross-Layer Index Reuse, https://arxiv.org/abs/2603.12201
- Frederic Lardinois, May 5th, 2026, The context window has been shattered: Subquadratic debuts a 12-million-token window: Subquadratic has launched a new AI architecture featuring a 12-million-token context window that outperforms GPT-5.5 on retrieval benchmarks, https://thenewstack.io/subquadratic-12-million-context-window/
- David Spuler, May 31st, 2026, Chapter 43. Long, Ultralong and Infinite Context, in book LLM Inference Optimization: State-of-the-Art Research, Table of Contents: https://www.aussieai.com/book/llm-inference-optimization https://www.amazon.com/dp/B0H3FKR39T
- Vishal Rajput, May 27, 2026, The Attention Problem Nobody Has Solved — Until Now? https://medium.com/aiguys/the-attention-problem-nobody-has-solved-until-now-bce5f3397ac9
- Subquadratic, May 5, 2026, How SSA Makes Long Context Practical, https://subq.ai/how-ssa-makes-long-context-practical
AI Books from Aussie AI
|
The Sweetest Lesson: Your Brain Versus AI: new book on AI intelligence theory:
Get your copy from Amazon: The Sweetest Lesson |
|
RAG Optimization: Accurate and Efficient LLM Applications:
new book on RAG architectures:
Get your copy from Amazon: RAG Optimization |
|
Generative AI Applications book:
Get your copy from Amazon: Generative AI Applications |
|
Generative AI programming book:
Get your copy from Amazon: Generative AI in C++ |
|
CUDA C++ Optimization book:
Get your copy from Amazon: CUDA C++ Optimization |
|
CUDA C++ Debugging book:
Get your copy from Amazon: CUDA C++ Debugging |
|
C++ AVX Optimization: CPU SIMD Vectorization:
Get your copy from Amazon: C++ AVX Optimization: CPU SIMD Vectorization |
|
C++ Ultra-Low Latency: Multithreading and Low-Level Optimizations:
Get your copy from Amazon: C++ Ultra-Low Latency |
More AI Research Topics
Read more about:
- 500+ LLM Inference Optimization Techniques
- What's Hot in LLM Inference Optimization in 2025?
- Inference Optimization Research
- « Research Home