Lossless Quantization
-
Last Updated 19 June, 2026
-
by David Spuler, Ph.D.
Lossless Quantization: Book Excerpts and Blog Articles
Free online book excerpts with full text chapters online and free PDF downloads, and the Aussie AI blog, including related articles:
- David Spuler, May 31st, 2026, Chapter 10. Quantization, in book LLM Inference Optimization: State-of-the-Art Research, Table of Contents: https://www.aussieai.com/book/llm-inference-optimization https://www.amazon.com/dp/B0H3FKR39T
Research on Lossless Quantization
Research papers include:
- Carl Franzen, May 20, 2026, Cohere cracks lossless quantization and native citations with first full Apache 2.0 licensed open model Command A+, https://venturebeat.com/technology/cohere-cracks-lossless-quantization-and-native-citations-with-first-full-apache-2-0-licensed-open-model-command-a
- Michael Helcig, Eldar Kurtic, Dan Alistarh, 4 May 2026, Statistically-Lossless Quantization of Large Language Models https://arxiv.org/abs/2605.02404 https://github.com/IST-DASLab/SLQ
- Moshik Hershcovitch, Andrew Wood, Leshem Choshen, Guy Girmonsky, Roy Leibovitz, Ilias Ennmouri, Michal Malka, Peter Chin, Swaminathan Sundararaman, Danny Harnik, 4 Jun 2025 (v2), ZipNN: Lossless Compression for AI Models, https://arxiv.org/abs/2411.05239
- Zeyu Yang, Tianyi Zhang, Jianwen Xie, Chuan Li, Zhaozhuo Xu, Anshumali Shrivastava, 3 Oct 2025, To Compress or Not? Pushing the Frontier of Lossless GenAI Model Weights Compression with Exponent Concentration, https://arxiv.org/abs/2510.02676
- Tianyi Zhang, Mohsen Hariri, Shaochen Zhong, Vipin Chaudhary, Yang Sui, Xia Hu, Anshumali Shrivastava, 1 Jan 2026 (v3), 70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11), https://arxiv.org/abs/2504.11651
- Hangming Zhang, Zheng Li, Qiang Yu, 20 Aug 2025, Quantization Meets Spikes: Lossless Conversion in the First Timestep via Polarity Multi-Spike Mapping, https://arxiv.org/abs/2508.14520
- Ke Yi, Jianwei Zhang, Zhiying Xu, Xinlong Yang, Yang Zhou, Minmin Sun, Zengke Liu, Tong Zhang, Junyang Lin, and Jingren Zhou, 2025, FPE2M2: Approaching Lossless and Efficient Quantization with Native Floating Point, Findings of the Association for Computational Linguistics: ACL 2025, pages 18350–18361, Vienna, Austria, Association for Computational Linguistics, https://aclanthology.org/2025.findings-acl.943/
- J. Wang, H. Liu, D. Feng, J. Ding and B. Ding, 2024, FP4-Quantization: Lossless 4bit Quantization for Large Language Models, 2024 IEEE International Conference on Joint Cloud Computing (JCC), Shanghai, China, 2024, pp. 61-67, doi: 10.1109/JCC62314.2024.00017, https://ieeexplore.ieee.org/document/10685437
- Mandar Karhade, MD. PhD., May 2026, Lossless Quantization: A Cohere Breakthrough: Cohere’s W4A4 lossless quantization could change economics of SOTA AI forever, https://medium.com/@AiDocTakes/lossless-quantization-a-cohere-breakthrough-afcf43f2a00d
- David Spuler, May 31st, 2026, Chapter 10. Quantization, in book LLM Inference Optimization: State-of-the-Art Research, Table of Contents: https://www.aussieai.com/book/llm-inference-optimization https://www.amazon.com/dp/B0H3FKR39T
More AI Research Topics
Read more about:
- 500+ LLM Inference Optimization Techniques
- What's Hot in LLM Inference Optimization in 2025?
- Inference Optimization Research
- « Research Home