AI Agent Compiles Triton Kernels Directly to PTX Without Triton Compiler
Summary
The paper investigates whether large language models can replace parts of the conventional compiler backend through a process the authors call AI lowering. An LLM agent translates Triton kernels directly into NVIDIA PTX, while an evaluation environment executes and verifies candidate code. Across twelve common kernels on Ada, Hopper, and Blackwell GPUs and ten kernels drawn from recent machine-learning papers, the approach reaches between 0.83x and 3.34x the performance of autotuned Triton. The largest reported gains come from transformations that the existing Triton lowering pipeline does not perform, including decoding packed binary weights directly into Tensor Core operands, assigning each thread a complete softmax row in tensor memory, and reusing overlapping convolution windows. The study reports 3.34x on BitDelta, 1.37x on FlashAttention, and up to 2.23x on convolution workloads. The authors attribute the results partly to a robust verification harness built on Volta, an existing PTX verifier. They substantially extend that verifier for modern architectures, including Blackwell's tcgen05 Tensor Core interface, managed tensor memory, descriptor-based operand layouts, and asynchronous execution coordinated by commits, waits, memory barriers, and proxy fences. The paper also discusses the difficulty of formalizing these features and the verifier's current limitations. Its conclusion is that AI-based compiler backends could eventually reduce the engineering effort needed to support new general-purpose and custom chips, although the reported results are an evaluation of a research system rather than a complete replacement for conventional compilers.