Study Finds Large Language Models Overpredict Social Punishment
Summary
Previous AI alignment work has largely focused on first-order social norms, such as recognizing that stealing is unacceptable. This study examines second-order expectations, or metanorms: judgments about who will enforce a norm, how they will respond, and when violators regulate themselves or observers regulate others. The researchers propose an evaluation framework with two dimensions, emotional appraisal and behavioral response, and introduce classification tasks for self-regulation and other-regulation. They release NormReact, a multi-perspective dataset containing 450 norm-violation scenarios hand-annotated for emotions and behavioral responses across the violator’s gender and the observer’s social closeness. Across six language models, the systems portray a harsher social world than humans do, overpredicting negative sanctions in situations where people would expect inaction. Agreement with human judgments also declines as the social distance between observers and violators increases. The authors argue that these errors could matter in norm-sensitive applications such as conflict mediation and policy simulation, where models may underrepresent tolerance, restraint, and relationship-dependent judgment while overrepresenting punishment.