Back to News
RSS feedarxiv.org

A Failure-Mode Framework for How AI Could Cause Civilizational Catastrophe

Summary

This arXiv article proposes a failure-mode framework for examining how advanced artificial intelligence could contribute to human extinction, irreversible civilizational collapse, or permanent human disempowerment. Its central argument is that catastrophic outcomes do not require an AI system to be conscious, hostile, or explicitly intended to harm people. Risks can instead emerge through four interacting pathways: autonomous misalignment, harmful human use, organizational failure, and competitive deployment. The severity of these pathways is shaped by capability, autonomy, external access, persistence, institutional safeguards, and whether people retain the capacity to recover from failures. The framework is deliberately non-operational and does not provide instructions for causing harm. Instead, it identifies causal conditions, intermediate quantities that could be studied empirically, and defensive research questions for understanding and reducing these risks.