Aussie AI
Kernel Synthesis
-
Last Updated 19 June, 2026
-
by David Spuler, Ph.D.
Kernel Synthesis: Book Excerpts and Blog Articles
Free online book excerpts with full text chapters online and free PDF downloads, and the Aussie AI blog, including related articles:
- David Spuler, May 31st, 2026, Chapter 55. Kernel Optimization Overview, in book LLM Inference Optimization: State-of-the-Art Research, Table of Contents: https://www.aussieai.com/book/llm-inference-optimization https://www.amazon.com/dp/B0H3FKR39T
Research on Kernel Synthesis
Research papers include:
- Sina Heidari, Dimitrios S. Nikolopoulos, 9 Apr 2026, FACT: Compositional Kernel Synthesis with a Three-Stage Agentic Workflow, https://arxiv.org/abs/2604.26666
- Weinan Dai, Hanlin Wu, Qiying Yu, Huan-ang Gao, Jiahao Li, Chengquan Jiang, Weiqiang Lou, Yufan Song, Hongli Yu, Jiaze Chen, Wei-Ying Ma, Ya-Qin Zhang, Jingjing Liu, Mingxuan Wang, Xin Liu, Hao Zhou, 27 Feb 2026, CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation, https://arxiv.org/abs/2602.24286
- Kris Shengjun Dong, Sahil Modi, Dima Nikiforov, Sana Damani, Edward Lin, Siva Kumar Sastry Hari, Christos Kozyrakis, 15 Feb 2026, KernelBlaster: Continual Cross-Task CUDA Optimization via Memory-Augmented In-Context Reinforcement Learning, https://arxiv.org/abs/2602.14293 https://github.com/NVlabs/KernelBlaster
- Charles Hong, Sahil Bhatia, Alvin Cheung, and Sophia Shao, 2025, Autocomp: LLM-Driven Code Optimization for Tensor Accelerators. In MLArchSys 2025, https://openreview.net/forum?id=bPdQZedlsr https://github.com/ucb-bar/autocomp
- Yoon Noh Lee, Yongseung Yu, and Yongjun Park. 2025. CUrator: An Efficient LLM Execution Engine with Optimized Integration of CUDA Libraries. In Proceedings of the 23rd ACM/IEEE International Symposium on Code Generation and Optimization (CGO '25). Association for Computing Machinery, New York, NY, USA, 209–224. https://doi.org/10.1145/3696443.3708944 https://dl.acm.org/doi/abs/10.1145/3696443.3708944
- Mingzhen Li, Hailong Yang, Shanjun Zhang, Fengwei Yu, Ruihao Gong, Yi Liu, Zhongzhi Luan, and Depei Qian. 2023. Exploiting Subgraph Similarities for Efficient Auto-tuning of Tensor Programs. In Proceedings of the 52nd International Conference on Parallel Processing (ICPP '23). Association for Computing Machinery, New York, NY, USA, 786–796. https://doi.org/10.1145/3605573.3605596 https://dl.acm.org/doi/10.1145/3605573.3605596
- Shiyang Li, Zijian Zhang, Winson Chen, Yuebo Luo, Mingyi Hong, Caiwen Ding, 3 Mar 2026, StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning, https://arxiv.org/abs/2603.02637
- Anne Ouyang, Simon Guo, Simran Arora, Alex L. Zhang, William Hu, Christopher Ré, Azalia Mirhoseini, 14 Feb 2025, KernelBench: Can LLMs Write Efficient GPU Kernels? https://arxiv.org/abs/2502.10517
- Anjiang Wei, Tianran Sun, Yogesh Seenichamy, Hang Song, Anne Ouyang, Azalia Mirhoseini, Ke Wang, Alex Aiken, 2 Dec 2025 (v2), Astra: A Multi-Agent System for GPU Kernel Performance Optimization, https://arxiv.org/abs/2509.07506 https://github.com/Anjiang-Wei/Astra
- Dezhi Ran, Shuxiao Xie, Mingfang Ji, Anmin Liu, Mengzhou Wu, Yuan Cao, Yuzhe Guo, Hao Yu, Linyi Li, Yitao Hu, Wei Yang, Tao Xie, 11 Feb 2026 (v2), KernelBand: Steering LLM-based Kernel Optimization via Hardware-Aware Multi-Armed Bandits, https://arxiv.org/abs/2511.18868
- Shiyi Cao, Ziming Mao, Joseph E. Gonzalez, Ion Stoica, 26 Feb 2026 (v2), K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model, https://arxiv.org/abs/2602.19128
- Haonan Li, Keyu Man, Partha Kanuparthy, Hanning Chen, Wei Sun, Sreen Tallam, Chenguang Zhu, Kevin Zhu, Zhiyun Qian, 14 Dec 2025 (v2), TritonForge: Profiling-Guided Framework for Automated Triton Kernel Optimization, https://arxiv.org/abs/2512.09196 https://github.com/zksainx/paper-notes/blob/main/docs/automated-kernel-generation/tritonforge.md
- AMD, May 2026 (accessed), Tensile documentation, https://rocm.docs.amd.com/projects/Tensile/en/latest/src/index.html
- Harrison Barclay, Ian Tramble, Michał Kukuła, Michel Migdal, Nicholai Tukanov and Leopold Cambier, Sep 02, 2025, Improving GEMM Kernel Auto-Tuning Efficiency on NVIDIA GPUs with Heuristics and CUTLASS 4.2, NVIDIA Technical Blog, https://developer.nvidia.com/blog/improving-gemm-kernel-auto-tuning-efficiency-on-nvidia-gpus-with-heuristics-and-cutlass-4-2/
- NVIDIA, May 2026 (accessed), NVIDIA Matmul Heuristics, https://docs.nvidia.com/cuda/nvidia-matmul-heuristics/
- Emergent Mind, 16 December 2025, Automated Triton Kernel Optimization, https://www.emergentmind.com/topics/automated-triton-kernel-optimization
- Tong Zheng, Haolin Liu, Chengsong Huang, Huiwen Bao, Sheng Zhang, Rui Liu, Runpeng Dai, Ruibo Chen, Chenxi Liu, Tianyi Xiong, Xidong Wu, Hongming Zhang, Heng Huang, 12 May 2026 (v2), LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling, https://arxiv.org/abs/2605.08083 https://github.com/zhengkid/AutoTTS
- David Spuler, May 31st, 2026, Chapter 55. Kernel Optimization Overview, in book LLM Inference Optimization: State-of-the-Art Research, Table of Contents: https://www.aussieai.com/book/llm-inference-optimization https://www.amazon.com/dp/B0H3FKR39T
AI Books from Aussie AI
|
The Sweetest Lesson: Your Brain Versus AI: new book on AI intelligence theory:
Get your copy from Amazon: The Sweetest Lesson |
|
RAG Optimization: Accurate and Efficient LLM Applications:
new book on RAG architectures:
Get your copy from Amazon: RAG Optimization |
|
Generative AI Applications book:
Get your copy from Amazon: Generative AI Applications |
|
Generative AI programming book:
Get your copy from Amazon: Generative AI in C++ |
|
CUDA C++ Optimization book:
Get your copy from Amazon: CUDA C++ Optimization |
|
CUDA C++ Debugging book:
Get your copy from Amazon: CUDA C++ Debugging |
|
C++ AVX Optimization: CPU SIMD Vectorization:
Get your copy from Amazon: C++ AVX Optimization: CPU SIMD Vectorization |
|
C++ Ultra-Low Latency: Multithreading and Low-Level Optimizations:
Get your copy from Amazon: C++ Ultra-Low Latency |
More AI Research Topics
Read more about:
- 500+ LLM Inference Optimization Techniques
- What's Hot in LLM Inference Optimization in 2025?
- Inference Optimization Research
- « Research Home