FP64 Emulation
-
Last Updated 12 June, 2026
-
by David Spuler, Ph.D.
Research on FP64 Emulation
Research papers include:
- Piotr Luszczek, Vijay Gadepally, LaToya Anderson, William Arcand, David Bestor, William Bergeron, Alex Bonn, Daniel J. Burrill, Chansup Byun, Michael Houle, Matthew Hubbell, Hayden Jananthan, Michael Jones, Peter Michaleas, Guillermo Morales, Julia Mullen, Andrew Prout, Albert Reuther, Antonio Rosa, Charles Yee, Jeremy Kepner, 28 Sep 2025, Performance and Numerical Aspects of Decompositional Factorizations with FP64 Floating-Point Emulation in INT8, https://arxiv.org/abs/2509.23565 https://ieee-hpec.org/wp-content/uploads/2025/09/127.pdf
- Rohail T., October 3, 2025, Decompositional Factorizations with FP64 Emulation in INT8 Demonstrate Performance and Numerical Profiles on Hopper GPUs, https://quantumzeitgeist.com/performance-decompositional-factorizations-fp64-emulation-int8-numerical-profiles-hopper/
- Shuntaro Ichimura, Takahiro Katagiri, Katsuhisa Ozaki, Takeshi Ogita, and Toru Nagai. Threaded accurate matrix-matrix multiplications with sparse matrix-vector multiplications. In 2018 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), page 1093–1102, 2018. https://ieeexplore.ieee.org/document/8425535
- Daichi Mukunoki, Katsuhisa Ozaki, Takeshi Ogita, and Toshiyuki Imamura. Accurate matrix multiplication on binary128 format accelerated by ozaki scheme. In Proceedings of the 50th International Conference on Parallel Processing (Lemont, IL, USA) (ICPP ’21), New York, NY, USA, 2021. Association for Computing Machinery. Article 78, 11 pages, https://dl.acm.org/doi/fullHtml/10.1145/3472456.3472493
- Daichi Mukunoki, Katsuhisa Ozaki, Takeshi Ogita, and Toshiyuki Imamura. DGEMM using tensor cores, and its accurate and reproducible versions. In Ponnuswamy Sadayappan, Bradford L. Chamberlain, Guido Juckeland, and Hatem Ltaief, editors, High Performance Computing, pages 230–248. Springer International Publishing, Cham, 2020. https://pmc.ncbi.nlm.nih.gov/articles/PMC7295351/ https://pmc.ncbi.nlm.nih.gov/articles/PMC7295351/pdf/978-3-030-50743-5_Chapter_12.pdf
- Daichi Mukunoki, Katsuhisa Ozaki, Takeshi Ogita, and Toshiyuki Imamura. Infinite-precision inner product and sparse matrix-vector multiplication using ozaki scheme with DOT2 on manycore processors. In International Conference on Parallel Processing and Applied Mathematics, page 40–54. Springer, 2022. https://link.springer.com/chapter/10.1007/978-3-031-30442-2_4
- Hiroyuki Ootomo, Katsuhisa Ozaki, and Rio Yokota. DGEMM on integer matrix multiplication unit. The International Journal of High Performance Computing Applications, 38(4):297–313, 2024. https://dl.acm.org/doi/abs/10.1177/10943420241239588
- Yuki Uchino, Katsuhisa Ozaki, Toshiyuki Imamura, 8 Aug 2025 (v2), High-Performance and Power-Efficient Emulation of Matrix Multiplication using INT8 Matrix Engines https://arxiv.org/abs/2508.03984
- Daichi Mukunoki, 25 Sep 2025 (v3), DGEMM without FP64 Arithmetic - Using FP64 Emulation and FP8 Tensor Cores with Ozaki Scheme, https://arxiv.org/abs/2508.00441
- K. Ozaki, T. Ogita, S. Oishi, and S. M. Rump. 2012. Error-free transformations of matrix multiplication by using fast routines of matrix multiplication and its applications. Numer. Algorithms 59, 1 (2012), 95–118. https://link.springer.com/article/10.1007/s11075-011-9478-1 (Original paper on DGEMM and FP64 emulation back in 2012.)
- Hiroyuki Ootomo, Katsuhisa Ozaki, Rio Yokota, 30 Mar 2024 (v4), DGEMM on Integer Matrix Multiplication Unit, https://arxiv.org/abs/2306.11975
- Cole Brower, Samuel Rodriguez Bernabeu, Jeff Hammond, John Gunnels, Sotiris S. Xanthea, Martin Ganahl, Andor Menczer, Örs Legeza, 6 Oct 2025, Mixed-precision ab initio tensor network state methods adapted for NVIDIA Blackwell technology via emulated FP64 arithmetic, https://arxiv.org/abs/2510.04795
- Katsuhisa Ozaki, Yuki Uchino, Toshiyuki Imamura, 27 Apr 2025 (v3), Ozaki Scheme II: A GEMM-oriented emulation of floating-point matrix multiplication using an integer modular technique, https://arxiv.org/abs/2504.08009
- Yuki Uchino, Katsuhisa Ozaki, Toshiyuki Imamura, 20 Sep 2024, Performance Enhancement of the Ozaki Scheme on Integer Matrix Multiplication Unit, https://arxiv.org/abs/2409.13313
- Samuel Rodriguez, July 2025, Floating Point Emulation in NVDIA Math Libraries: Optimizing Floating Point Precision, CERN, July 1-2, 2025, Geneve, Switzerland, https://indico.cern.ch/event/1538409/contributions/6521976/attachments/3096181/5485165/cern-talk.pdf
- Y. Uchino, Q. Ma, T. Imamura, K. Ozaki and P. L. Gutsche, "Emulation of Complex Matrix Multiplication based on the Chinese Remainder Theorem," ISC High Performance 2026 Research Paper Proceedings (41st International Conference), Hamburg, Germany, 2026, pp. 1-12, doi: 10.23919/ISC.2026.11520500, https://ieeexplore.ieee.org/abstract/document/11520500 https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=11520500
- Angelika Schwarz, Anton Anders, Cole Brower, Harun Bayraktar, John Gunnels, Kate Clark, RuQing G. Xu, Samuel Rodriguez, Sebastien Cayrols, Pawel Tabaszewski, and Victor Podlozhnyuk, 2026, Guaranteed DGEMM Accuracy While Using Reduced Precision Tensor Cores Through Extensions of the Ozaki Scheme. In Proceedings of the Supercomputing Asia and International Conference on High Performance Computing in Asia Pacific Region (SCA/HPCAsia '26). Association for Computing Machinery, New York, NY, USA, 91–101, https://doi.org/10.1145/3773656.3773670 https://dl.acm.org/doi/full/10.1145/3773656.3773670
- Yuki Uchino, Katsuhisa Ozaki, Toshiyuki Imamura, 6 Apr 2026 (v2), Double-Precision Matrix Multiplication Emulation via Ozaki-II Scheme with FP8 Quantization, https://arxiv.org/abs/2603.10634
- J. Dongarra, J. Gunnels, H. Bayraktar, A. Haidar and D. Ernst, 2025, Accelerating Supercomputing: AI-Hardware-Driven Innovation for Speed and Efficiency, 2025 IEEE High Performance Extreme Computing Conference (HPEC), Wakefield, MA, USA, 2025, pp. 1-7, doi: 10.1109/HPEC67600.2025.11196413, https://ieeexplore.ieee.org/document/11196413
- Devangi N. Parikh, Robert A. van de Geijn, Greg M. Henry, 30 Dec 2024 (v2), Cascading GEMM: High Precision from Low Precision, https://arxiv.org/abs/2303.04353
- Hiroyuki Ootomo and Rio Yokota, 2022, Recovering single precision accuracy from Tensor Cores while surpassing the FP32 theoretical peak performance. Int. J. High Perform. Comput. Appl. 36, 4 (Jul 2022), 475–491, https://doi.org/10.1177/10943420221090256 https://dl.acm.org/doi/abs/10.1177/10943420221090256
Ozaki FP64 Emulation
- Shuntaro Ichimura, Takahiro Katagiri, Katsuhisa Ozaki, Takeshi Ogita, and Toru Nagai. Threaded accurate matrix-matrix multiplications with sparse matrix-vector multiplications. In 2018 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), page 1093–1102, 2018. https://ieeexplore.ieee.org/document/8425535
- Daichi Mukunoki, Katsuhisa Ozaki, Takeshi Ogita, and Toshiyuki Imamura. Accurate matrix multiplication on binary128 format accelerated by ozaki scheme. In Proceedings of the 50th International Conference on Parallel Processing (Lemont, IL, USA) (ICPP ’21), New York, NY, USA, 2021. Association for Computing Machinery. Article 78, 11 pages, https://dl.acm.org/doi/fullHtml/10.1145/3472456.3472493
- Daichi Mukunoki, Katsuhisa Ozaki, Takeshi Ogita, and Toshiyuki Imamura. DGEMM using tensor cores, and its accurate and reproducible versions. In Ponnuswamy Sadayappan, Bradford L. Chamberlain, Guido Juckeland, and Hatem Ltaief, editors, High Performance Computing, pages 230–248. Springer International Publishing, Cham, 2020. https://pmc.ncbi.nlm.nih.gov/articles/PMC7295351/ https://pmc.ncbi.nlm.nih.gov/articles/PMC7295351/pdf/978-3-030-50743-5_Chapter_12.pdf
- Daichi Mukunoki, Katsuhisa Ozaki, Takeshi Ogita, and Toshiyuki Imamura. Infinite-precision inner product and sparse matrix-vector multiplication using ozaki scheme with DOT2 on manycore processors. In International Conference on Parallel Processing and Applied Mathematics, page 40–54. Springer, 2022. https://link.springer.com/chapter/10.1007/978-3-031-30442-2_4
- Hiroyuki Ootomo, Katsuhisa Ozaki, and Rio Yokota. DGEMM on integer matrix multiplication unit. The International Journal of High Performance Computing Applications, 38(4):297–313, 2024. https://dl.acm.org/doi/abs/10.1177/10943420241239588
- Yuki Uchino, Katsuhisa Ozaki, Toshiyuki Imamura, 8 Aug 2025 (v2), High-Performance and Power-Efficient Emulation of Matrix Multiplication using INT8 Matrix Engines https://arxiv.org/abs/2508.03984
- Daichi Mukunoki, 25 Sep 2025 (v3), DGEMM without FP64 Arithmetic - Using FP64 Emulation and FP8 Tensor Cores with Ozaki Scheme, https://arxiv.org/abs/2508.00441
- K. Ozaki, T. Ogita, S. Oishi, and S. M. Rump. 2012. Error-free transformations of matrix multiplication by using fast routines of matrix multiplication and its applications. Numer. Algorithms 59, 1 (2012), 95–118. https://link.springer.com/article/10.1007/s11075-011-9478-1 (Original paper on DGEMM and FP64 emulation back in 2012.)
- Hiroyuki Ootomo, Katsuhisa Ozaki, Rio Yokota, 30 Mar 2024 (v4), DGEMM on Integer Matrix Multiplication Unit, https://arxiv.org/abs/2306.11975
- Katsuhisa Ozaki, Yuki Uchino, Toshiyuki Imamura, 27 Apr 2025 (v3), Ozaki Scheme II: A GEMM-oriented emulation of floating-point matrix multiplication using an integer modular technique, https://arxiv.org/abs/2504.08009
- Yuki Uchino, Katsuhisa Ozaki, Toshiyuki Imamura, 20 Sep 2024, Performance Enhancement of the Ozaki Scheme on Integer Matrix Multiplication Unit, https://arxiv.org/abs/2409.13313
- Samuel Rodriguez, July 2025, Floating Point Emulation in NVDIA Math Libraries: Optimizing Floating Point Precision, CERN, July 1-2, 2025, Geneve, Switzerland, https://indico.cern.ch/event/1538409/contributions/6521976/attachments/3096181/5485165/cern-talk.pdf
More AI Research Topics
Read more about:
- 500+ LLM Inference Optimization Techniques
- What's Hot in LLM Inference Optimization in 2025?
- Inference Optimization Research
- « Research Home