LLMs Will Call AI Compilers, Not Replace Them
Summary
This opinion article argues that the claim that large language models will replace AI compilers depends on what “replace” means. If an LLM must perform compilation inside its weights, producing IR, assembly, or binaries through token generation, the approach would be inefficient: memory planning, tiling, scheduling, and instruction selection can be handled by specialized algorithms far more cheaply. It would also be difficult to verify, because a nondeterministic and opaque translator could produce different outputs and make numerical regressions hard to reproduce. The more practical architecture is tool-based orchestration. An LLM could inspect an intermediate representation, select compiler passes, run autotuners and profilers, compare results with a baseline, and iterate, while the compiler retains the search loops, cost models, memory planners, and verification machinery. The article argues that accumulated tuning records can make the model’s reasoning a one-time investment, provided hints are tied to the model, shapes, hardware, and compiler version. The author also distinguishes orchestration from authorship: LLMs are more likely to change how kernels, MLIR lowerings, backends, and compiler heuristics are written. Generated code still needs reference-based correctness checks, real-hardware performance tests, numerical or equivalence checking, and strong evaluation harnesses, since weak tests can reward code that is fast but wrong. In this view, LLMs become capable authors working inside compiler infrastructure rather than replacements for it. AI compiler engineering shifts toward clean interfaces, verifiable intermediate representations, fast search, and defining what correctness means.