Aussie AI
LLM Inference Optimization: State-of-the-Art Research
-
by David Spuler
Book Blurb
For research engineers, scientists and academics looking for cutting-edge new ideas and state-of-the-art algorithms for LLM inference optimization. An extensive survey of LLM Inference optimization techniques with research paper lists and guidance about what's new, what's widely used in industry, and what's coming soon from research labs.
Highlights
- 500+ LLM inference optimization techniques categorized.
- 100+ detailed chapters on all of our favorite ones.
- 50+ breakthroughs: new, emerging, and future research.
Table of Contents
- Part III: Quantization
10. Quantization
11. Low-Bit Quantization
12. Block-Scaled Quantization (BSQ)
13. Fixed-Point and Block-Floating Point
14. Bitshift Quantization
15. Weight Clustering
16. Rare Quantization Techniques
Part IV: Compute Optimizations
17. Prefill Optimizations
18. Serving Disaggregation
19. Mixture-of-Experts (MoE)
20. MoE Attention
21. FFN Optimizations
22. FFN Fusion
23. FFN Folding
24. Flash FFN
Part V: Transformer Components
25. Tokenizer and Vocabulary
26. Unembedding Module
27. Model Loader
28. Encoders & Decoders
29. Activation Functions
30. Softmax
31. Normalization
32. Positional Encoding
33. Embedding and Unembedding
VI: Attention Kernels
34. Attention Overview
35. Flash Attention
36. Paged Attention
37. Radix Attention
38. MLA Attention
39. QKV Compute Optimizations
40. DeepSeek Sparse Attention (DSA)
41. Other Attention Kernels
42. Hybrid Attention
43. Long, Ultralong and Infinite Context
Part VII: Decoding Algorithms
44. Decoding Algorithms
45. Speculative Decoding Introduction
46. Optimizing Speculative Decoding
47. Model-Based Speculative Decoding
48. Parallel Drafting
49. Eagle, Medusa, and FR-Spec
50. Lookup Decoding
51. Generalized Speculative Decoding
52. Structured Decoding
53. Aggressive Decoding
54. Edit Decoding
Part VIII: Kernel Optimization
55. Kernel Optimization Overview
56. Integer Bitwise Operations
57. Floating-Point Bit Tricks
58. Kernel Fusion Overview
59. Kernel Fusion Examples
60. CUDA Optimizations
61. Grace CPU Optimizations
62. Hopper & Blackwell GPUs
63. Rubin and Feynman GPUs
64. Vectorization
65. Branchless Coding
Part IX: Matrix Multiplication Kernels
66. MatMul/GEMM Introduction
67. MatMul Research
68. Strassen and Winograd Matrix Multiplication
Part X: Caching
69. Caching Optimizations
70. KV Caching Overview
71. RAG Caching
72. KV Cache Compression
73. KV Pruning and Fusion
74. Mooncake
75. KV Cache Data Correction
76. Prefix KV Caching
77. Non-Prefix KV Caching
Part XI: Pruning
78. Pruning
79. Unstructured Pruning
80. Structured Pruning
81. Depth Pruning
82. Layer Pruning
83. Early Exit
84. Layer Skipping
85. Shallow Decoder & Shallow Prefill
86. Layer Reordering
87. Width Pruning
88. Length Pruning
89. Prompt and Context Compression
90. Token Pruning
91. Zero Skipping
92. Weight Sharing
Part XII: Data Structures & Algorithms
93. Radix Trees
94. Longest Common Substring
95. Vector Dot Product
96. Top-K Vector Algorithm
97. Vector Norms
98. Vector Databases
99. Tensors
100. Lookup Tables & Precomputation
101. Conditional Computation
102. Weight Precomputation
Part XIII: Advanced Research
103. Advanced Number Systems
104. Knowledge Distillation
105. Zero-Multiplication Models
106. Additive Models
107. Arithmetic Optimization Research
108. Ensemble Multi-Model Architectures
109. Logarithmic Models
110. Neural Architecture Search
Part XIV: Reasoning Model Optimizations
111. Reasoning Inference Optimization
112. Chain-of-Thought Token Reduction
113. Small Reasoning Models
114. Concept Token Reasoning
Appendix 1. List of 500 Techniques
- Index
Preface
About the Author
About the Contributors
Table of Contents
Part I: Introduction to LLM Inference Optimization
1. Speed is the Need
2. Introduction to LLM Inference Optimization
3. SOTA Inference Stacks
4. Accuracy Versus Speed
Part II: Top-Level Architecture Optimizations
5. Harness Engineering
6. Deployment Architecture
7. RAG Architectures
8. Agentic Architecture Optimizations
9. Adaptive Inference
|
LLM Inference Optimization Book: • Online: Table of Contents • Buy: LLM Inference Optimization |
|
New book: LLM Inference Optimization:
Get your copy from Amazon: LLM Inference Optimization |