Aussie AI

LLM Inference Optimization: State-of-the-Art Research

  • by David Spuler

Book Blurb

For research engineers, scientists and academics looking for cutting-edge new ideas and state-of-the-art algorithms for LLM inference optimization. An extensive survey of LLM Inference optimization techniques with research paper lists and guidance about what's new, what's widely used in industry, and what's coming soon from research labs.

Highlights

  • 500+ LLM inference optimization techniques categorized.
  • 100+ detailed chapters on all of our favorite ones.
  • 50+ breakthroughs: new, emerging, and future research.

Table of Contents

    Preface

    About the Author

    About the Contributors

    Table of Contents

    Part I: Introduction to LLM Inference Optimization

        1. Speed is the Need

        2. Introduction to LLM Inference Optimization

        3. SOTA Inference Stacks

        4. Accuracy Versus Speed

    Part II: Top-Level Architecture Optimizations

        5. Harness Engineering

        6. Deployment Architecture

        7. RAG Architectures

        8. Agentic Architecture Optimizations

        9. Adaptive Inference

  • Part III: Quantization

        10. Quantization

        11. Low-Bit Quantization

        12. Block-Scaled Quantization (BSQ)

        13. Fixed-Point and Block-Floating Point

        14. Bitshift Quantization

        15. Weight Clustering

        16. Rare Quantization Techniques

    Part IV: Compute Optimizations

        17. Prefill Optimizations

        18. Serving Disaggregation

        19. Mixture-of-Experts (MoE)

        20. MoE Attention

        21. FFN Optimizations

        22. FFN Fusion

        23. FFN Folding

        24. Flash FFN

    Part V: Transformer Components

        25. Tokenizer and Vocabulary

        26. Unembedding Module

        27. Model Loader

        28. Encoders & Decoders

        29. Activation Functions

        30. Softmax

        31. Normalization

        32. Positional Encoding

        33. Embedding and Unembedding

    VI: Attention Kernels

        34. Attention Overview

        35. Flash Attention

        36. Paged Attention

        37. Radix Attention

        38. MLA Attention

        39. QKV Compute Optimizations

        40. DeepSeek Sparse Attention (DSA)

        41. Other Attention Kernels

        42. Hybrid Attention

        43. Long, Ultralong and Infinite Context

    Part VII: Decoding Algorithms

        44. Decoding Algorithms

        45. Speculative Decoding Introduction

        46. Optimizing Speculative Decoding

        47. Model-Based Speculative Decoding

        48. Parallel Drafting

        49. Eagle, Medusa, and FR-Spec

        50. Lookup Decoding

        51. Generalized Speculative Decoding

        52. Structured Decoding

        53. Aggressive Decoding

        54. Edit Decoding

    Part VIII: Kernel Optimization

        55. Kernel Optimization Overview

        56. Integer Bitwise Operations

        57. Floating-Point Bit Tricks

        58. Kernel Fusion Overview

        59. Kernel Fusion Examples

        60. CUDA Optimizations

        61. Grace CPU Optimizations

        62. Hopper & Blackwell GPUs

        63. Rubin and Feynman GPUs

        64. Vectorization

        65. Branchless Coding

    Part IX: Matrix Multiplication Kernels

        66. MatMul/GEMM Introduction

        67. MatMul Research

        68. Strassen and Winograd Matrix Multiplication

    Part X: Caching

        69. Caching Optimizations

        70. KV Caching Overview

        71. RAG Caching

        72. KV Cache Compression

        73. KV Pruning and Fusion

        74. Mooncake

        75. KV Cache Data Correction

        76. Prefix KV Caching

        77. Non-Prefix KV Caching

    Part XI: Pruning

        78. Pruning

        79. Unstructured Pruning

        80. Structured Pruning

        81. Depth Pruning

        82. Layer Pruning

        83. Early Exit

        84. Layer Skipping

        85. Shallow Decoder & Shallow Prefill

        86. Layer Reordering

        87. Width Pruning

        88. Length Pruning

        89. Prompt and Context Compression

        90. Token Pruning

        91. Zero Skipping

        92. Weight Sharing

    Part XII: Data Structures & Algorithms

        93. Radix Trees

        94. Longest Common Substring

        95. Vector Dot Product

        96. Top-K Vector Algorithm

        97. Vector Norms

        98. Vector Databases

        99. Tensors

        100. Lookup Tables & Precomputation

        101. Conditional Computation

        102. Weight Precomputation

    Part XIII: Advanced Research

        103. Advanced Number Systems

        104. Knowledge Distillation

        105. Zero-Multiplication Models

        106. Additive Models

        107. Arithmetic Optimization Research

        108. Ensemble Multi-Model Architectures

        109. Logarithmic Models

        110. Neural Architecture Search

    Part XIV: Reasoning Model Optimizations

        111. Reasoning Inference Optimization

        112. Chain-of-Thought Token Reduction

        113. Small Reasoning Models

        114. Concept Token Reasoning

    Appendix 1. List of 500 Techniques

  • Index

 

LLM Inference Optimization Book:

Online: Table of Contents

Buy: LLM Inference Optimization

LLM Inference Optimization New book: LLM Inference Optimization:
  • 50+ research breakthroughs: current, new and emerging
  • 100+ chapters on inference efficiency techniques
  • 500+ LLM inference optimization techniques
  • State-of-the-art research & literature review

Get your copy from Amazon: LLM Inference Optimization