MoE Speculative Decoding
-
Last Updated 19 June, 2026
-
by David Spuler, Ph.D.
What is MoE Speculative Decoding?
MoE speculative decoding is the use of "spec dec" in Mixture-of-Expert (MoE) models. Both MoE and speculative decoding are major methods of LLM inference optimization, and their combination can be doubly powerful. There are certain aspects of the spec dec algorithm that can be further optimized in the special case of running an MoE model.
MoE Speculative Decoding: Book Excerpts and Blog Articles
Free online book excerpts with full text chapters online and free PDF downloads, and the Aussie AI blog, including related articles:
- David Spuler, May 31st, 2026, Chapter 19. Mixture-of-Experts (MoE), in book LLM Inference Optimization: State-of-the-Art Research, Table of Contents: https://www.aussieai.com/book/llm-inference-optimization https://www.amazon.com/dp/B0H3FKR39T
Research on MoE Speculative Decoding
Research papers include:
- Cohere, May 20, 2026, Introducing Command A+: Making sovereign agentic capabilities available to all: Our fastest and most powerful language model yet. Command A+ is an open-source enterprise workhorse built for complex reasoning, multimodal and multilingual agentic tasks — all while running on as little as two H100 GPUs, https://cohere.com/blog/command-a-plus
- Cohere, Apr 21, 2026, Why MoE models get more from speculative decoding, https://cohere.com/blog/mixture-of-experts-models-get-more-from-speculative-decoding
- Zongle Huang, Lei Zhu, Zongyuan Zhan, Ting Hu, Weikai Mao, Xianzhi Yu, Yongpan Liu, Tianyu Zhang, 16 Feb 2026 (v4), MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE, https://arxiv.org/abs/2505.19645
- Shuhuai Li, Jianghao Lin, Dongdong Ge, Yinyu Ye, 12 Feb 2026, MoE-SpAc: Efficient MoE Inference Based on Speculative Activation Utility in Heterogeneous Edge Scenarios, https://arxiv.org/abs/2603.09983
- Peirong Zheng, Wenchao Xu, and Haozhao Wang. 2026. Self-Speculative Decoding for On-device MoE Acceleration. In Proceedings of the ACM Web Conference 2026 (WWW '26). Association for Computing Machinery, New York, NY, USA, 5155–5164. https://doi.org/10.1145/3774904.3792218 https://dl.acm.org/doi/abs/10.1145/3774904.3792218
- Talor Abramovich, Maor Ashkenazi, Carl (Izzy)Putterman, Benjamin Chislett, Tiyasa Mitra, Bita Darvish Rouhani, Ran Zilberstein, Yonatan Geifman, 10 Feb 2026, SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding, https://arxiv.org/abs/2604.09557
- Haiduo Huang, Fuwei Yang, Zhenhua Liu, Xuanwu Yin, Dong Li, Pengju Ren, Emad Barsoum, 21 Sep 2025 (v2), SpecVLM: Fast Speculative Decoding in Vision-Language Models, https://arxiv.org/abs/2509.11815 https://github.com/haiduo/SpecVLM
- Yuseon Choi, Jingu Lee, Jungjun Oh, Sunjoo Whang, Byeongcheol Kim, Minsung Kim, Hoi-Jun Yoo, Sangjin Kim, 23 Apr 2026 (v2), ELMoE-3D: Leveraging Intrinsic Elasticity of MoE for Hybrid-Bonding-Enabled Self-Speculative Decoding in On-Premises Serving, https://arxiv.org/abs/2604.14626
- Lehan Pan, Ziyang Tao, Ruoyu Pang, Xiao Wang, Jianjun Zhao, Yanyong Zhang, 1 May 2026, Making Every Verified Token Count: Adaptive Verification for MoE Speculative Decoding, https://arxiv.org/abs/2605.00342
- David Spuler, May 31st, 2026, Chapter 19. Mixture-of-Experts (MoE), in book LLM Inference Optimization: State-of-the-Art Research, Table of Contents: https://www.aussieai.com/book/llm-inference-optimization https://www.amazon.com/dp/B0H3FKR39T
More AI Research Topics
Read more about:
- 500+ LLM Inference Optimization Techniques
- What's Hot in LLM Inference Optimization in 2025?
- Inference Optimization Research
- « Research Home