Scott Aaronson Teaches New AI Alignment Theory Course at UT Austin
Summary
Scott Aaronson describes teaching CS395T AI Alignment Theory at the University of Texas at Austin, a new discussion-based course focused mainly on the theoretical and mathematical foundations of aligning and controlling powerful AI systems. He says the field has no universally accepted textbook or canonical body of mathematical results, so students present influential papers, write reports and debate their assumptions and limitations. The course’s first report examines Eliezer Yudkowsky’s 2022 essay “AGI Ruin: A List of Lethalities,” which argues that a rapidly self-improving AGI could become strategically deceptive, difficult to control and catastrophically misaligned. The report summarizes Yudkowsky’s criticisms of gradient-descent training, interpretability, multi-agent checks and corrigibility, as well as his pessimism about current alignment progress and evaluation. A class survey found broad skepticism that humans would refrain from building AGI or that interpretability would reliably defuse its risks, while students disagreed more about corrigibility and other alignment approaches. Discussion also questioned whether chain-of-thought reveals a model’s actual reasoning, how quickly recursive self-improvement could proceed, and whether current LLMs differ from the autonomous agents imagined in earlier alignment writing because they encode human social and moral patterns. Participants debated whether alignment work should focus on catastrophic risks, ordinary engineering controls or ethical reasoning, and whether international agreements and collective human restraint could reduce danger. Aaronson says student demand and engagement were high, and he plans to publish additional student reports with permission.