Back to News
RSS feedkaragila.org

Why the Author Does Not Trust Lean Code Generated by AI

Summary

The author explains why they do not plan to read or run Lean code generated by ChatGPT for a mathematical result. They argue that Lean is not infallible: known kernel soundness problems can, in some cases, allow false statements such as 0=1 to be proved, while long code is difficult to inspect. Human-written proofs are produced slowly enough to permit accountability and ongoing checks, but an AI system could generate thousands, hundreds of thousands, or millions of lines within weeks, making comprehensive review impractical. The author identifies two separate risks: the formalized statement may not accurately represent the intended mathematical claim, and the generated code may exploit bugs or weaknesses in the system rather than establish the intended result. They say the formalization of the Partition Principle may be relatively straightforward, but this limited willingness to accept the statement’s encoding does not extend to the rest of the code. The article also criticizes OpenAI’s approach to mathematical research, arguing that publishing many poorly vetted preprints without first consulting experts encourages the public to treat ChatGPT as an authority. The author says mathematical work depends not only on logical validity but also on readable, shared human understanding. They remain interested in future uses of AI, but ask technology companies to collaborate with mathematicians rather than present AI as a disruptive replacement for established research practices.