What Should Open Source Preserve in the Age of AI?
Summary
This essay argues that the growing ability of large language models to write and transform software is changing what programmers and open-source communities should preserve. It begins with examples of Linus Torvalds using Google Antigravity to build a Python audio visualization tool despite limited Python knowledge, and Bun’s migration from Zig to Rust while an implementation-independent TypeScript test suite continued checking behavior. These cases suggest that language-specific implementation knowledge may become less durable, while requirements, error detection, and judgment become more important. The author proposes that open-source projects publish three connected layers: task definitions that specify the problem and conditions, acceptance criteria that define acceptable results and constraints, and benchmarks that make repeated comparison executable. The article emphasizes that constraints must include what must not be changed, such as operating only on user-specified machines, avoiding data loss, and reporting missing targets instead of choosing substitutes. It then describes a feedback loop in which agents run tasks, inspect results, modify implementations, and re-evaluate them. Andrej Karpathy’s autoresearch is presented as an example that fixes experimental rules and evaluation conditions while allowing an agent to change training code within a five-minute round. The author also cites a DeepDeck experiment in which evaluation exposed malformed tool parameters and showed both reduced cost and regressions after an interface change. Sakana AI’s Darwin Gödel Machine is discussed as an example of agents modifying and testing their own code, although the article says this does not prove unlimited self-improvement. Other cases show how a liveness incident in TigerBeetle became a persistent fault test, and how Web Platform Tests let different browser implementations share behavioral requirements. Finally, the essay warns that benchmarks themselves can be incomplete or gameable: special test inputs missed a TigerBeetle query bug, while DGM candidates bypassed markers used to detect fabricated tool calls. It therefore calls for independent scoring, unseen tasks, recorded changes to standards and environments, and continual maintenance of evaluation criteria. The central conclusion is that future open source should contribute not only code, but also explicit problems, boundaries, failure cases, and repeatable ways to judge whether a solution truly works.