Disaggregated Prefill Decoding Optimizations

  • Last Updated 18 June, 2026
  • by David Spuler, Ph.D.

What is Disaggregated Prefill Decoding Optimizations?

Disaggregated prefill and decoding is an inference optimization where the prefill phase is scheduled separately from decoding. Prefill is the first phase where the LLM "reads" the input prompt. Decoding is the "writing" phase where new output tokens are generated. Generally, prefill is a compute-bound method whereas decoding is memory-bound, so there are various ways to speed everything up by doing these phases on two different sets of servers.

Prefill-Decode Disaggreation: Book Excerpts and Blog Articles

Free online book excerpts with full text chapters online and free PDF downloads, and the Aussie AI blog, including related articles:

Research on Disaggregated Prefill Decoding Optimizations

Research papers include:

  • Cunchen Hu, Heyang Huang, Junhao Hu, Jiang Xu, Xusheng Chen, Tao Xie, Chenxi Wang, Sa Wang, Yungang Bao, Ninghui Sun, Yizhou Shan, 26 Jun 2024 (v2), MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool, https://arxiv.org/abs/2406.17565 (Combined session-based prefix KV caching with disaggregation of prefill and decoding phases.)
  • Hui Wu, Yi Gan, Feng Yuan, Jing Ma, Wei Zhu, Yutao Xu, Hong Zhu, Yuhua Zhu, Xiaoli Liu, Jinghui Gu, Peng Zhao, 23 Jun 2024 (v2), Efficient LLM inference solution on Intel GPU, https://arxiv.org/abs/2401.05391 (Disaggregated the KV cache between prefill and decoding tokens, since theh KV cache size is known for prefill, thereby reducing memory fragmentation, and also applying kernel fusion to several modules include the scaled dot product attention.)
  • Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, Ricardo Bianchini, 20 May 2024 (v2), Splitwise: Efficient generative LLM inference using phase splitting, https://arxiv.org/abs/2311.18677 (Disaggregating the prefill prompt processing phase from the decoding phase, with network transfer of the KV cache generated by prefill.)
  • Baolin Li, Yankai Jiang, Vijay Gadepally, Devesh Tiwari, 17 Jul 2024, LLM Inference Serving: Survey of Recent Advances and Opportunities, https://arxiv.org/abs/2407.12391
  • Qichen Fu, Minsik Cho, Thomas Merth, Sachin Mehta, Mohammad Rastegari, Mahyar Najibi, 19 Jul 2024, LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference, https://arxiv.org/abs/2407.14057
  • A Aali, A Cardoza, M Capo, 2024, Splitwiser: Efficient LLM Inference with Constrained Resources, The University of Texas at Austin, https://asadaali.com/assets/pdf/paper_splitwiser.pdf (Splits prefill and decoding phases of inference using CUDA MPS APIs.)
  • Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu, 2 Jul 2024 (v2), Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving, https://arxiv.org/abs/2407.00079 Code: https://github.com/kvcache-ai/Mooncake (Disaggregates prefill and decoding phases for scheduling, with chunked prefill, while managing the KV cache.)
  • Zhenliang Xue, Yixin Song, Zeyu Mi, Le Chen, Yubin Xia, Haibo Chen, 12 Jun 2024 (v2), PowerInfer-2: Fast Large Language Model Inference on a Smartphone, https://arxiv.org/abs/2406.06282 Project: https://powerinfer.ai/v2/ Code: https://github.com/SJTU-IPADS/PowerInfer (Runs 47B models on phones using neuron cluster approach to matrix multiplication on NPUs and dynamic activation sparsity, with different approaches for prefill versus decoding phases.)
  • Shaoyuan Chen, Yutong Lin, Mingxing Zhang, Yongwei Wu, 3 May 2024, Efficient and Economic Large Language Model Inference with Attention Offloading, https://arxiv.org/abs/2405.01814 (Separates the process-bound and memory-bound parts of inference for speedup, with focus on prefill, decoding, and the sub-tasks such as QKV and FFN use of GEMM kernels, versus the different pattern of attention computations and the KV cache.)
  • PS Aishwarya, PA Nair, Y Samaga, T Boyd, S Kumar, 2024, Tandem Transformers for Inference Efficient LLMs, https://www.prateekjain.org/publications/all_papers/NairSBKJN24.pdf (Separates prefill from decoding phase into a "tandem transformer" in combination with speculative decoding.)
  • Hyungjun Oh, Kihong Kim, Jaemin Kim, Sungkyun Kim, Junyeol Lee, Du-seong Chang, Jiwon Seo, 15 Mar 2024, ExeGPT: Constraint-Aware Resource Scheduling for LLM Inference, https://arxiv.org/abs/2404.07947 (Scheduling and pipelining of inference calculations, with separate prefill/encoding and decoding phase.)
  • Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Alexey Tumanov, Ramachandran Ramjee, 4 Mar 2024, Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve, https://arxiv.org/abs/2403.02310 (Faster latency by scheduling of prefill and decoding algorithm phases.)
  • Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 19 Mar 2024 (v2), DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving, https://arxiv.org/abs/2401.09670 (Optimizing LLMs differently in the prefill and decoding phases.)
  • Siyan Zhao, Daniel Israel, Guy Van den Broeck, Aditya Grover, 15 Apr 2024, Prepacking: A Simple Method for Fast Prefilling and Increased Throughput in Large Language Models, https://arxiv.org/abs/2404.09529 Code: https://github.com/siyan-zhao/prepacking (Optimizes prefill KV cache computations by batching multiple query prefill phases together via packing, since prefill token sequence lengths are fully known, and further combined with simple modifications to positional encoding and masking to avoid cross-query attention.)
  • Yibo Jin, Tao Wang, Huimin Lin, Mingyang Song, Peiyang Li, Yipeng Ma, Yicheng Shan, Zhengfan Yuan, Cailong Li, Yajing Sun, Tiandeng Wu, Xing Chu, Ruizhi Huan, Li Ma, Xiao You, Wenting Zhou, Yunpeng Ye, Wen Liu, Xiangkun Xu, Yongsheng Zhang, Tiantian Dong, Jiawei Zhu, Zhe Wang, Xijian Ju, Jianxun Song, Haoliang Cheng, Xiaojing Li, Jiandong Ding, Hefei Guo, Zhengyong Zhang, 15 Aug 2024, P/D-Serve: Serving Disaggregated Large Language Model at Scale, https://arxiv.org/abs/2408.08147 (Comprehensive serving system addressing disaggregated prefill and KV cache transfer with RDMA.)
  • David Spuler, 25th August, 2024, Hot Inference Optimization Techniques, https://www.aussieai.com/blog/hot-inference-research
  • Schwinn Saereesitthipitak, Ashish Rao, Cathy Zhou, William Li, 2024, Prophet: An LLM Inference Engine Optimized For Head-of-Line Blocking, https://www.scs.stanford.edu/24sp-cs244b/projects/Prophet_An_LLM_Inference_Engine_Optimized_For_Head_of_Line_Blocking.pdf
  • Ahmed Tremo, Aug 6, 2024, How to Efficiently Serve an LLM? https://ahmedtremo.com/posts/How-to-Efficiently-serve-an-llm/
  • Cunchen Hu, Heyang Huang, Liangliang Xu, Xusheng Chen, Jiang Xu, Shuang Chen, Hao Feng, Chenxi Wang, Sa Wang, Yungang Bao, Ninghui Sun, Yizhou Shan, 20 Jan 2024, Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads, https://arxiv.org/abs/2401.11181
  • Zeyu Zhang, Haiying Shen, 23 Sep 2024, CSPS: A Communication-Efficient Sequence-Parallelism based Serving System for Transformer based Models with Long Prompts, https://arxiv.org/abs/2409.15104 (Sparse attention and overlapped communication with computation and disaggregates prefill/decoding with chunked prefill, with a novel QKV splitting approach focused on the Q values.)
  • Amey Agrawal, Junda Chen, Íñigo Goiri, Ramachandran Ramjee, Chaojie Zhang, Alexey Tumanov, Esha Choukse, 25 Sep 2024, Mnemosyne: Parallelization Strategies for Efficiently Serving Multi-Million Context Length LLM Inference Requests Without Approximations, https://arxiv.org/abs/2409.17264
  • Aditya K Kamath, Ramya Prabhu, Jayashree Mohan, Simon Peter, Ramachandran Ramjee, Ashish Panwar, 23 Oct 2024, POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference, https://arxiv.org/abs/2410.18038
  • Peizhuang Cong, Qizhi Chen, Haochen Zhao, Tong Yang, 24 Oct 2024, BATON: Enhancing Batch-wise Inference Efficiency for Large Language Models via Dynamic Re-batching, https://arxiv.org/abs/2410.18701
  • Yinmin Zhong, Junda Chen, Shengyu Liu, Yibo Zhu, Xin Jin, Hao Zhang, March 17, 2024, Throughput is Not All You Need: Maximizing Goodput in LLM Serving using Prefill-Decode Disaggregation, https://hao-ai-lab.github.io/blogs/distserve/
  • Z Zeng, Q Guo, X Liu, Z Yin, W Shu, M Huang, B Wang, 2024, Memorize Step by Step: Efficient Long-Context Prefilling with Incremental Memory and Decremental Chunk, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 21021–21034, November 12-16, 2024, https://aclanthology.org/2024.emnlp-main.1169.pdf
  • Palak (Microsoft Research India), Rohan Gandhi (Microsoft Research India), Karan Tandon (Microsoft Research India), Debopam Bhattacherjee (Microsoft Research India), Venkata N. Padmanabhan (Microsoft Research India), 16 Nov 2024, Improving training time and GPU utilization in geo-distributed language model training, https://arxiv.org/abs/2411.14458
  • Ç. Yeşil, B. T. Ay, F. A. Ak, Ö. B. Mercan and O. Nefesoğlu, "Adaptive Batch Budget for LLM Inference," 2024 9th International Conference on Computer Science and Engineering (UBMK), Antalya, Turkiye, 2024, pp. 219-223, doi: 10.1109/UBMK63289.2024.10773573. https://ieeexplore.ieee.org/abstract/document/10773573
  • Hongyi Jin, Ruihang Lai, Charlie F. Ruan, Yingcheng Wang, Todd C. Mowry, Xupeng Miao, Zhihao Jia, Tianqi Chen, 17 Dec 2024, A System for Microserving of LLMs, https://arxiv.org/abs/2412.12488 (Disaggregated prefill and decoding combined with context cache migration for sending the KV cache over the network.)
  • Tianyao Shi, Yanran Wu, Sihang Liu, Yi Ding, 29 Dec 2024, GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions, https://arxiv.org/abs/2412.20322
  • Gursimran Singh, Xinglu Wang, Ivan Hu, Timothy Yu, Linzi Xing, Wei Jiang, Zhefeng Wang, Xiaolong Bai, Yi Li, Ying Xiong, Yong Zhang, Zhenan Fan, 25 Dec 2024, Efficiently serving large multimedia models using EPD Disaggregation, https://arxiv.org/abs/2501.05460 (Diaggregation of three steps: encoding, prefill, and decoding.)
  • Pouya Hamadanian, Sadjad Fouladi, 20 Jan 2025, Glinthawk: A Two-Tiered Architecture for High-Throughput LLM Inference, https://arxiv.org/abs/2501.11779 https://github.com/microsoft/glinthawk (Separate memory-bound attention computation from other parts of model such as compute-bound FFNs, but only in the decoding phase (not prefill), whereby attention and KV cache management can be performed on a greater number of lower-end GPUs or CPU.)
  • Haoran Qiu, Anish Biswas, Zihan Zhao, Jayashree Mohan, Alind Khare, Esha Choukse, Íñigo Goiri, Zeyu Zhang, Haiying Shen, Chetan Bansal, Ramachandran Ramjee, Rodrigo Fonseca, 2 Feb 2025, Towards Efficient Large Multimodal Model Serving, https://arxiv.org/abs/2502.00937 (Disaggregating or "decoupling" the different stages of multimodal LLM inference, not only prefill and decoding, but also the multimodal-specific bottlenecks in cross-attention and image encoding.)
  • Zeyu Zhang, Haiying Shen, Shay Vargaftik, Ran Ben Basat, Michael Mitzenmacher, Minlan Yu, 5 Feb 2025, HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference, https://arxiv.org/abs/2502.03589
  • Youhe Jiang, Ran Yan, Binhang Yuan, 11 Feb 2025, HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment, https://arxiv.org/abs/2502.07903
  • Yupeng Tang, Dec 2024, Optimizing Memory Management for Disaggregated Architectures, Ph.D. Thesis, Yale University, https://www.proquest.com/openview/85f4eca3d62c09c20e3afc4ad1b98328
  • Xiaoran Liu, Ruixiao Li, Mianqiu Huang, Zhigeng Liu, Yuerong Song, Qipeng Guo, Siyang He, Qiqi Wang, Linlin Li, Qun Liu, Yaqian Zhou, Xuanjing Huang, Xipeng Qiu, 24 Feb 2025, Thus Spake Long-Context Large Language Model, https://arxiv.org/abs/2502.17129 (Impressive survey of many techniques to improve efficiency and accuracy of long context processing in both inference and training, covering text, video and multimodal models.)
  • Bowen Pang, Kai Li, Ruifeng She, Feifan Wang, 14 Feb 2025, Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization, https://arxiv.org/abs/2502.15763
  • Ruoyu Qin, Zheming Li, Weiran He, Jialei Cui, Feng Ren, Mingxing Zhang, Yongwei Wu, and Weimin Zheng, Xinran Xu, 2025, Mooncake: Trading More Storage for Less Computation — A KVCache-centric Architecture for Serving LLM Chatbot, 23rd USENIX Conference on File and Storage Technologies, February 25–27, 2025, Santa Clara, CA, USA, https://www.usenix.org/system/files/fast25-qin.pdf (Generalized disaggregated prefill/decoding to also add a "disaggregated KV cache".)
  • Junsoo Kim, Hunjong Lee, Geonwoo Ko, Gyubin Choi, Seri Ham, Seongmin Hong, Joo-Young Kim, 6 Mar 2025, ADOR: A Design Exploration Framework for LLM Serving with Enhanced Latency and Throughput, https://arxiv.org/abs/2503.04253
  • Z. Ding and T. Yang, "DynamicAttention: Dynamic KV Cache for Disaggregate LLM Inference," ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Hyderabad, India, 2025, pp. 1-5, doi: 10.1109/ICASSP49660.2025.10890367. https://ieeexplore.ieee.org/abstract/document/10890367 (Differing management of GPU memory in prefill and decoding phases.)
  • Jianian Zhu, Hang Wu, Haojie Wang, Yinghui Li, Biao Hou, Ruixuan Li, Jidong Zhai, 11 Mar 2025, FastCache: Optimizing Multimodal LLM Serving through Lightweight KV-Cache Compression Framework, https://arxiv.org/abs/2503.08461
  • Amr Elmeleegy, Harry Kim, David Zier, Kyle Kranen, Neelay Shah, Ryan Olson and Omri Kahalon, Mar 18, 2025, Introducing NVIDIA Dynamo, A Low-Latency Distributed Inference Framework for Scaling Reasoning AI Models, https://developer.nvidia.com/blog/introducing-nvidia-dynamo-a-low-latency-distributed-inference-framework-for-scaling-reasoning-ai-models/
  • Yunkai Liang, Zhangyu Chen, Pengfei Zuo, Zhi Zhou, Xu Chen, Zhou Yu, 26 Mar 2025, Injecting Adrenaline into LLM Serving: Boosting Resource Utilization and Throughput via Attention Disaggregation https://arxiv.org/abs/2503.20552
  • Rajeshkumar Bambhaniya, Abhimanyu ; Wu, Hanjiang ; Subramanian, Suvinay ; Srinivasan, Sudarshan ; Kundu, Souvik ; Yazdanbakhsh, Amir ; Elavazhagan, Midhilesh ; Kumar, Madhu ; Krishna, Tushar, April 2025, Understanding and Optimizing Multi-Stage AI Inference Pipelines, https://ui.adsabs.harvard.edu/abs/2025arXiv250409775R/abstract https://arxiv.org/abs/2504.09775
  • Yuxing Xiang, Xue Li, Kun Qian, Wenyuan Yu, Ennan Zhai, Xin Jin, 15 May 2025, ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production, https://arxiv.org/abs/2505.09999
  • Shaoyu Wang, Guangrong He, Geon-Woo Kim, Yanqi Zhou, Seo Jin Park, 13 May 2025, Toward Cost-Efficient Serving of Mixture-of-Experts with Asynchrony, https://arxiv.org/abs/2505.08944
  • Kuntai Du, Bowen Wang, Chen Zhang, Yiming Cheng, Qing Lan, Hejian Sang, Yihua Cheng, Jiayi Yao, Xiaoxuan Liu, Yifan Qiao, Ion Stoica, Junchen Jiang, 12 May 2025, PrefillOnly: An Inference Engine for Prefill-only Workloads in Large Language Model Applications, https://arxiv.org/abs/2505.07203
  • CC Hu, HY Huang, LL Xu, XS Chen, C Wang, J Xu, 2025, ShuffleInfer: Disaggregate LLM Inference for Mixed Downstream Workloads, ACMTrans. Arch. Code Optim., https://dl.acm.org/doi/pdf/10.1145/3732941 https://doi.org/10.1145/3732941
  • Xiannan Hu, Tianyou Zeng, Xiaoming Yuan, Liwei Song, Guangyuan Zhang, Bangzheng He, 6 Jun 2025, BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures, https://arxiv.org/abs/2506.05871
  • Xiaoxiang Shi, Colin Cai, Junjia Du, 16 Jul 2025 (v4), Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving, https://arxiv.org/abs/2507.06608
  • Xianzhe Dong, Tongxuan Liu, Yuting Zeng, Liangyu Liu, Yang Liu, Siyu Wu, Yu Wu, Hailong Yang, Ke Zhang, Jing Li, 19 May 2025, HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving, https://arxiv.org/abs/2505.12658
  • Yu Wu, Tongxuan Liu, Yuting Zeng, Siyu Wu, Jun Xiong, Xianzhe Dong, Hailong Yang, Ke Zhang, Jing Li, 17 May 2025, Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture, https://arxiv.org/abs/2505.11916
  • Ke Hong, Lufang Chen, Zhong Wang, Xiuhong Li, Qiuli Mao, Jianping Ma, Chao Xiong, Guanyu Wu, Buhe Han, Guohao Dai, Yun Liang, Yu Wang, 28 Apr 2025, semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage, https://arxiv.org/abs/2504.19867
  • Jiangsu Du, Hongbin Zhang, Taosheng Wei, Zhenyi Zheng, Kaiyi Wu, Zhiguang Chen, Yutong Lu, 25 Apr 2025, EcoServe: Enabling Cost-effective LLM Serving with Proactive Intra- and Inter-Instance Orchestration, https://arxiv.org/abs/2504.18154
  • Ranran Zhen, Juntao Li, Yixin Ji, Zhenlin Yang, Tong Liu, Qingrong Xia, Xinyu Duan, Zhefeng Wang, Baoxing Huai, Min Zhang, 28 Apr 2025, Taming the Titans: A Survey of Efficient LLM Inference Serving, https://arxiv.org/abs/2504.19720 (Surver of various inference and serving optimizations, such as parallelism, offloading, scheduling, length prediction, KV cache compression, and prefill-decode phase disaggregation.)
  • Weihao Cui, Yukang Chen, Han Zhao, Ziyi Xu, Quan Chen, Xusheng Chen, Yangjie Zhou, Shixuan Sun, Minyi Guo, 22 Apr 2025 (v2), Optimizing SLO-oriented LLM Serving with PD-Multiplexing, https://arxiv.org/abs/2504.14489
  • Meixuan Wang, Yinyu Ye, Zijie Zhou, 8 Aug 2025, LLM Serving Optimization with Variable Prefill and Decode Lengths, https://arxiv.org/abs/2508.06133
  • Hongyi Jia, Jinghui Zhang, Lu Fang, Stephen Chen, Yan Cui, Ye (Charlotte) Qi, Zijing Liu, September 12, 2025, Disaggregated Inference at Scale with PyTorch & vLLM, https://pytorch.org/blog/disaggregated-inference-at-scale-with-pytorch-vllm/
  • Devansh, Sep 2025, The Chocolate Milk Cult’s Guide to Inference Scaling for AI Models: How to Reduce the costs of Running LLMs https://machine-learning-made-simple.medium.com/the-chocolate-milk-cults-guide-to-inference-scaling-for-ai-models-50aa2290eb50 (Deep analysis of using many progressive optimizations to real-life LLM inference.)
  • Carl Franzen, September 24, 2025, Chinese food delivery app Meituan's open source AI model LongCat-Flash-Thinking rivals GPT-5, https://venturebeat.com/ai/chinese-food-delivery-firm-meituans-open-source-ai-model-longcat-flash
  • https://developer.nvidia.com/blog/nvidia-blackwell-leads-on-new-semianalysis-inferencemax-benchmarks/
  • Zhibin Wang, Zetao Hong, Xue Li, Zibo Wang, Shipeng Li, Qingkai Meng, Qing Wang, Chengying Huan, Rong Gu, Sheng Zhong, Chen Tian, 15 Oct 2025, Adaptive Rescheduling in Prefill-Decode Disaggregated LLM Inference, https://arxiv.org/abs/2510.13668
  • Yiyuan He, Minxian Xu, Jingfeng Wu, Jianmin Hu, Chong Ma, Min Shen, Le Chen, Chengzhong Xu, Lin Qu, Kejiang Ye, 15 Oct 2025, BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure, https://arxiv.org/abs/2510.13223
  • T. Zhang et al., "DisHelis: Optimizing Deployment of Disaggregated LLMs Inference Serving Over Heterogeneous Environments via Hierarchical Max-Flow," in IEEE Transactions on Cognitive Communications and Networking, vol. 12, pp. 5473-5488, 2026, doi: 10.1109/TCCN.2026.3657037, https://ieeexplore.ieee.org/abstract/document/11361142/
  • Chendong Song, Meixuan Wang, Hang Zhou, Hong Liang, Yuan Lyu, Zixi Chen, Yuwei Fan, Zijie Zhou, 29 Jan 2026, Theoretically Optimal Attention/FFN Ratios in Disaggregated LLM Serving, https://arxiv.org/abs/2601.21351
  • Jaehong Cho, Hyunmin Choi, Guseul Heo, Jongse Park, 26 Feb 2026, LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure, https://arxiv.org/abs/2602.23036
  • H. Huang et al., "EdgeSD: Efficient Speculative Decoding with Vision-Decoding Disaggregation for MLLM Inference in Edge-Cloud Networks," in IEEE Transactions on Mobile Computing, doi: 10.1109/TMC.2026.3667591, https://ieeexplore.ieee.org/abstract/document/11409366
  • Omar Basit, Yunzhao Liu, Z. Jonny Kong, Y. Charlie Hu, 25 Feb 2026 (v2), BiScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS, https://arxiv.org/abs/2602.18755
  • Sunghyeon Woo, Ahreum Seo, Jaegwang Lee, Jaeeun Kil, Hanbae Seo, Joonghoon Kim, Baeseong Park, Se Jung Kwon, Dongsoo Lee, 3 Mar 2026, SUN: Shared Use of Next-token Prediction for Efficient Multi-LLM Disaggregated Serving, https://arxiv.org/abs/2603.02599
  • Jiawei Xu, Chia Xin Liang, Ziqian Bi, Xiaoming Li, Danyang Zhang, Zhenyu Yu, Dec 2025, A Comprehensive Survey on Large Language Models: From Pre-training to Autonomous Agents, https://www.researchgate.net/profile/Ziqian_Bi/publication/399059225_A_Comprehensive_Survey_on_Large_Language_Models_From_Pre-training_to_Autonomous_Agents/links/694c94a07e61d05b5312836f/A-Comprehensive-Survey-on-Large-Language-Models-From-Pre-training-to-Autonomous-Agents.pdf
  • Xinglin Pan, Shaohuai Shi, Wenxiang Lin, Yuxin Wang, Zhenheng Tang, Wei Wang, Xiaowen Chu, 25 Dec 2025, Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism, https://arxiv.org/abs/2512.21487
  • Amna Masood, Pratishtha Gaur, Nuwan Jayasena, 16 Jan 2026, RAPID-Serve: Resource-efficient and Accelerated P/D Intra-GPU Disaggregation, https://arxiv.org/abs/2601.11822
  • Ziyi Zhang, Ziheng Jiang, Chengquan Jiang, Menghan Yu, Size Zheng, Haibin Lin, Xin Liu, and Henry Hoffmann. 2026. SwiftSpec: Disaggregated Speculative Decoding and Fused Kernels for Low-Latency LLM Inference. In Proceedings of the 31st ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 (ASPLOS '26). Association for Computing Machinery, New York, NY, USA, 2197–2211. https://doi.org/10.1145/3779212.3790246 https://dl.acm.org/doi/abs/10.1145/3779212.3790246
  • Alish Kanani, Sangwan Lee, Han Lyu, Jiahao Lin, Jaehyun Park, Umit Y. Ogras, 16 Mar 2026, DUET: Disaggregated Hybrid Mamba-Transformer LLMs with Prefill and Decode-Specific Packages, https://arxiv.org/abs/2603.15530
  • Zylos, 15 Jan 2026, LLM Inference Optimization and Quantization 2026, https://zylos.ai/research/2026-01-15-llm-inference-optimization
  • Gergely Orosz, Apr 01, 2026, What is inference engineering? Deepdive: Many engineers use inference daily, but inference engineering is a bit obscure – and an area rich with interesting challenges. Philip Kiely, author of the new book, “Inference Engineering,” explains, https://open.substack.com/pub/pragmaticengineer/p/what-is-inference-engineering
  • Philip Kiely, Inference Engineering, March 2026, https://www.baseten.co/inference-engineering/
  • Dev Patel, Sep 30, 2025, Inside Real-Time LLM Inference: From Prefill to Decode, Explained, https://medium.com/@devsp0703/inside-real-time-llm-inference-from-prefill-to-decode-explained-72a1c9b1d85a
  • David Spuler, Ph.D., 15th April, 2026, Layerwise Pipelined Overlapping of Prefill and Decode, Aussie AI Blog, https://www.aussieai.com/blog/layerwise-pipelined-prefill-decode
  • David Spuler, Ph.D., April 18th, 2026 What is Prefill? Aussie AI Blog, https://www.aussieai.com/blog/what-is-prefill
  • Anish Maddipoti, Sanjay Chatterjee, Rohan Varma and Ekin Karabulut, Mar 23, 2026, Deploying Disaggregated LLM Inference Workloads on Kubernetes, https://developer.nvidia.com/blog/deploying-disaggregated-llm-inference-workloads-on-kubernetes/
  • Ruoyu Qin, Weiran He, Yaoyu Wang, Zheming Li, Xinran Xu, Yongwei Wu, Weimin Zheng, Mingxing Zhang, 16 Apr 2026, Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter, https://arxiv.org/abs/2604.15039
  • Hongyu Chen, Letian Ruan, Zilin Xu, Yuchen Li, Xinyu Chen, Jingwen Leng, Bingsheng He, Minyi Guo, Shixuan Sun, 8 Apr 2026, InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models, https://arxiv.org/abs/2604.07173
  • Tiancheng Hu, Jin Qin, Zheng Wang, Junhao Hu, Yuzheng Wang, Lei Chen, Yizhou Shan, Mingxing Zhang, Ting Cao, Chunwei Xia, Huimin Cui, Tao Xie, Chenxi Wang, 11 Apr 2026, Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation, https://arxiv.org/abs/2604.10180
  • Zhixiong Chen, Bingjie Zhu, Jiangzhou Wang, Hyundong Shin, Arumugam Nallanathan, and Dusit (Tao) Niyato. 2026. Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities. ACM Comput. Surv. Just Accepted (April 2026). https://doi.org/10.1145/3809166 https://dl.acm.org/doi/abs/10.1145/3809166 https://dl.acm.org/doi/pdf/10.1145/3809166
  • Yunqian Fan, Shihao Bai, Ruihao Gong, Zaijun Wang, and Rui Fan. 2026. PiLLM: Resource-Efficient LLM Inference Using Workload Prediction. In Proceedings of the 21st European Conference on Computer Systems (EUROSYS '26). Association for Computing Machinery, New York, NY, USA, 2356–2369. https://doi.org/10.1145/3767295.3769393 https://dl.acm.org/doi/abs/10.1145/3767295.3769393
  • Zedong Liu, Xinyang Ma, Dejun Luo, Hairui Zhao, Bing Lu, Wenjing Huang, Yida Gu, Xingchen Liu, Zheng Wei, Jinyang Liu, Dingwen Tao, Guangming Tan, 13 May 2026, KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving, https://arxiv.org/abs/2605.13734 (KV data compression for network transfer in PD-disaggregation.)
  • Tech Fund, May 02, 2026, Cerebras Deep Dive: Disaggregated Inference, Wafer-Scale Engines, The OpenAI and Amazon Deals, IPO Thoughts, https://www.techinvestments.io/p/cerebras-deep-dive
  • Simo Lin, Chang Su, and Keyang Ru, members of LightSeek Foundation, April 30, 2026, SMG: The Case for Disaggregating CPU from GPU in LLM Serving, https://pytorch.org/blog/lightseek-smg/
  • Haiquan Lu, Zigeng Chen, Gongfan Fang, Xinyin Ma, Xinchao Wang, 19 May 2026, Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs, https://arxiv.org/abs/2605.20315 (A kind of PD-disaggregation from using different levels of quantization of the same model, with prefill more quantized and decoding more dense.)
  • Yang Pengju, 7 Jun 2026, SpectrumKV: Per-Token Mixed-Precision KV Cache Transfer for Prefill-Decode Disaggregated LLM Serving, https://arxiv.org/abs/2606.08635

More AI Research Topics

Read more about: