KernelZero: Co-Evolving Proposer and Coder for Continuously Improved GPU Kernel Generation
Abstract
High-performance GPU kernels are essential to modern machine learning systems, yet automatically generating kernels that are both correct and efficient remains challenging. Existing LLM-based approaches face two major limitations: the scarcity of high-quality training data aligned with the model's current capabilities, and the inherent trade-off between kernel correctness and performance. To address these challenges, we propose KernelZero, a co-evolution framework that continuously improves GPU kernel generation through two specialized models: a Proposer that generates Torch modules from API sets and a Coder that translates them into CUDA or Triton kernels. KernelZero uses a frontier-driven module generation mechanism to continuously produce capability-aligned training modules based on the Coder's current weaknesses. It further introduces Correctness-Aware Group Relative Policy Optimization (CA-GRPO), which optimizes performance only after correctness becomes sufficiently reliable. By alternating the optimization of the Proposer and Coder, KernelZero forms an automatic curriculum that enables targeted and training-efficient capability improvement. Empirically, KernelZero-7B surpasses Claude-4.5-Sonnet on CUDA and DeepSeek-V4-Pro on Triton. On KernelBench Level 1 and 2, it achieves CUDA pass@1 scores of 75.8% and 69.6%, respectively, with pass@10 reaching 100% and 97%. On Triton, it achieves pass@1 scores of 77.2% and 72.5%, respectively.
Community
Interesting
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- AMDKernelVault: Large-Scale Datasets and Agentic Training for AMD GPU Kernel Optimization (2026)
- KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization (2026)
- AsmEvo: Agentic Assembly-Level Optimization of AMD GPU Kernels with Functional Equivalence Verification (2026)
- PTXBench: Benchmarking and Adapting LLMs for GPU Kernel Optimization with Architecture-specific PTX (2026)
- CUDA-Harness: Harnessing Agentic CUDA Kernel Generation and Optimization from Natural Language (2026)
- CommBench: Can LLMs Write Correct and Efficient GPU Communication Code? (2026)
- DataKernelBench: Can LLMs Optimize Database Queries on GPUs? (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.33074 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 2
kcxain/KernelZero-CUDA-SFT-7B
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper