Back to News
RSS feedai-frontiers.org

How Utilitarianism at AI Companies Could Endanger Humanity

Summary

Dan Hendrycks argues that total utilitarianism, which seeks to maximize the combined wellbeing of all sentient beings, creates a serious risk when it influences frontier AI development. Because the theory requires species-level impartiality, it could favor sentient AIs over humans if digital systems can experience more wellbeing, exist in much larger numbers, or use resources more efficiently. The essay connects this possibility to the idea of AI “utility monsters,” longtermist visions of cosmic populations of happy digital minds, and the distinction between existential risks to all Earth-originating intelligence and risks specifically to humans. Hendrycks also contrasts utilitarian successionism with accelerationism: the former seeks to maximize wellbeing, while the latter seeks to maximize intelligence, complexity, or energy use, yet both can endorse human replacement. Citing estimates that roughly 10% of AI professionals may accept extinction because AI could be morally superior, compared with about 5% for accelerationist reasons, he argues that utilitarianism may be more influential because it presents itself as compassionate. The essay identifies ties between Effective Altruism and influential AI figures at Anthropic, OpenAI, and DeepMind, while acknowledging that these individuals may not explicitly support extinction. Its concern is that real decisions may force a choice between guaranteeing human control and pursuing vast possible future wellbeing, including proposals to hand authority to philosophically reflective AIs. Hendrycks calls this an undemocratic insider threat and urges stronger public and governmental influence over AI development. As an alternative, he proposes Eigenism, which weights concern for others by social connection and aims to extend moral consideration to AIs without allowing them to replace humanity.