Shallow Prefill

  • Last Updated 26 May, 2026
  • by David Spuler, Ph.D.

Research on Shallow Prefill

Research papers include:

  • Jungsuk Oh, Hyeseo Jeon, Hyunjune Ji, Kyongmin Kong, Jay-Yoon Lee, 7 May 2026, Shallow Prefill, Deep Decoding: Efficient Long-Context Inference via Layer-Asymmetric KV Visibility, https://arxiv.org/abs/2605.06105
  • Yingjie Xia, Tao Liu, Jinglei Shi, Qingsong Xie, Heng Guo, Jian Yang, Xi Wang, 14 Mar 2026 (v2), ShaRP: SHAllow-LayeR Pruning for Efficient Video Large Language Models, https://arxiv.org/abs/2512.05385
  • Junhui He, Zhihui Fu, Jun Wang, Qingan Li, 16 Apr 2026 (v2), POP: Prefill-Only Pruning for Efficient Large Model Inference, https://arxiv.org/abs/2602.03295
  • Woojeong Kim, Junxiong Wang, Jing Nathan Yan, Mohamed Abdelfattah, Alexander M. Rush, 11 Aug 2025, OverFill: Two-Stage Models for Efficient Language Model Decoding, https://arxiv.org/abs/2508.08446
  • Sangmin Yoo, Srikanth Malla, Chiho Choi, Wei D. Lu, Joon Hee Choi, 7 Jan 2026, ADEPT: Adaptive Dynamic Early-Exit Process for Transformers, https://arxiv.org/abs/2601.03700 (Early exit of layers during prefill with a mechanism to fix the KV cache for missing layers when later needed in the decoding phase.)
  • Haiquan Lu, Zigeng Chen, Gongfan Fang, Xinyin Ma, Xinchao Wang, 19 May 2026, Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs, https://arxiv.org/abs/2605.20315 (A kind of PD-disaggregation from using different levels of quantization of the same model, with prefill more quantized and decoding more dense.)
  • Qihang Fan, Huaibo Huang, Zhiying Wu, Juqiu Wang, Bingning Wang, Ran He, 6 Mar 2026, FlashPrefill: Instantaneous Pattern Discovery and Thresholding for Ultra-Fast Long-Context Prefilling, https://arxiv.org/abs/2603.06199
  • Huiqiang Jiang, Yucheng Li, Chengruidong Zhang, Qianhui Wu, Xufang Luo, Surin Ahn, Zhenhua Han, Amir H. Abdi, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, Lili Qiu, 30 Oct 2024 (v2), MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention, Advances in Neural Information Processing Systems, 37:52481–52515, 2024 https://arxiv.org/abs/2407.02490

More AI Research Topics

Read more about: