Model Loading
-
Last Updated 19 June, 2026
-
by David Spuler, Ph.D.
Model Loading: Book Excerpts and Blog Articles
Free online book excerpts with full text chapters online and free PDF downloads, and the Aussie AI blog, including related articles:
- David Spuler, March 2024, Chapter 19. Encoders & Decoders, in book "Generative AI in C++", https://www.aussieai.com/book/ch19-intro-encoder-decoder
- David Spuler, March 2024, Generative AI in C++: Coding Transformers and LLMs, https://www.aussieai.com/book/toc PDF: https://www.aussieai.com/pdf/BOOK-Generative-AI-CPP-Spuler-2024.pdf
- David Spuler, May 31st, 2026, Chapter 27. Model Loader, in book LLM Inference Optimization: State-of-the-Art Research, Table of Contents: https://www.aussieai.com/book/llm-inference-optimization https://www.amazon.com/dp/B0H3FKR39T
Research on Model Loading
Research papers include:
- Shuming Shi, Enbo Zhao, Deng Cai, Leyang Cui, Xinting Huang, Huayang Li, 16 Jan 2024, Inferflow: an Efficient and Highly Configurable Inference Engine for Large Language Models, https://arxiv.org/abs/2401.08294 Source: https://github.com/inferflow/inferflow
- Jing Liu, Ruihao Gong, Mingyang Zhang, Yefei He, Jianfei Cai, Bohan Zhuang, 13 Jun 2024, ME-Switch: A Memory-Efficient Expert Switching Framework for Large Language Models, https://arxiv.org/abs/2406.09041 (How to load multiple experts for MoE in a memory-efficient way using mixed-precision quantization based on identifying the few salient channels that need higher precision, as an alternative to multi-LoRA.)
- Xueyuan Han, Zinuo Cai, Yichu Zhang, Chongxin Fan, Junhan Liu, Ruhui Ma, Rajkumar Buyya, 9 Sep 2024 (v2), Hermes: Memory-Efficient Pipeline Inference for Large Models on Edge Devices, https://arxiv.org/abs/2409.04249 (Pipelining of model layer-wise loading and inference for memory-efficient inference.)
- Yifan Sui, Hao Wang, Hanfei Yu, Yitao Hu, Jianxun Li, Hao Wang, 20 May 2025, ServerlessLoRA: Minimizing Latency and Cost in Serverless Inference for LoRA-Based LLMs, https://arxiv.org/abs/2505.14468
- David Spuler, March 2024, Chapter 19. Encoders & Decoders, in book "Generative AI in C++", https://www.aussieai.com/book/ch19-intro-encoder-decoder
- David Spuler, March 2024, Generative AI in C++: Coding Transformers and LLMs, https://www.aussieai.com/book/toc PDF: https://www.aussieai.com/pdf/BOOK-Generative-AI-CPP-Spuler-2024.pdf
- T. P. Patel, "Fast Warm-Start Techniques for Large Language Model Inference," 2026 IEEE 5th International Multidisciplinary Conference on Engineering Technology (IMCET), Beirut, Lebanon, 2026, pp. 1-8, doi: 10.1109/IMCET69180.2026.11503728, https://ieeexplore.ieee.org/abstract/document/11503728
- Charles Frye, Jonathan Belotti, Erik Bernhardsson, Akshat Bubna, May 12, 2026, Cutting inference cold starts by 40x with LP, FUSE, C/R, and cuda-checkpoint, https://modal.com/blog/truly-serverless-gpus (Spinning up a new instance with GPUs in seconds.)
- David Spuler, May 31st, 2026, Chapter 27. Model Loader, in book LLM Inference Optimization: State-of-the-Art Research, Table of Contents: https://www.aussieai.com/book/llm-inference-optimization https://www.amazon.com/dp/B0H3FKR39T
More AI Research Topics
Read more about:
- 500+ LLM Inference Optimization Techniques
- What's Hot in LLM Inference Optimization in 2025?
- Inference Optimization Research
- « Research Home