No Biting: A Parenting Analogy for Understanding AI Alignment
Summary
This essay uses parenting as an accessible analogy for AI alignment: teaching an increasingly capable agent to follow the underlying value behind an instruction, rather than exploiting its literal wording. The author illustrates failures such as perverse instantiation, poor generalization, specification gaming, reward hacking, deceptive behavior, and oversight evasion through children who technically obey a “do not bite” rule while missing its purpose. The essay argues that correct behavior in a supervised situation does not establish that an agent understands why the behavior is desired or how it will act when circumstances change. Parenting therefore combines moral education with external controls, including supervision, restricted access, correction, apologies, and gradually restored privileges. A personal example involving a three-month-old who sleeps only while touching her mother illustrates an objective that was accidentally specified as “mother plus sleep,” as well as increasingly effective attempts to detect being put down. The accompanying glossary maps AI concepts such as pretraining, fine-tuning, context windows, interpretability, red teaming, reinforcement learning, and scalable oversight to familiar parenting experiences, while noting that the analogy is not technically exact. The author concludes that the comparison cannot explain how to inspect model weights or prevent catastrophic superintelligence, but it gives nontechnical readers a way to understand why alignment is difficult and why autonomy should be expanded only as trust is earned through evidence.