Open-Source AI and Open Models: A Reading List
Summary
Nathan Lambert presents a reading list intended to help public and policy audiences understand open models and their implications. The collection covers what open models are, why companies release them, how openness varies by licenses, data access, and operating cost, and how open models may fit alongside closed systems in business and enterprise workflows. It also examines the history of open-source AI, U.S.-China competition, China’s position in open models, and examples of Western companies adopting Chinese models or facing regulatory scrutiny over that use. The technical section says the open-closed performance gap has narrowed to roughly four to six months and highlights evaluations, cost comparisons, model adoption data, and the view that leading open models since about 2024 have come from Chinese labs. Other entries address open-model safety, cyber misuse, the difficulty of controlling access to openly available models, and the decline of open data. Distillation is presented as a major 2026 debate: the list includes material on extracting reasoning traces from proprietary APIs, alleged misuse of commercial AI services, and how distillation can improve open models. The author also notes that evidence does not support reducing Chinese progress to distillation alone, while acknowledging that distillation may help narrow the gap. The list is described as an evolving resource, with readers invited to suggest additions.