Back to News
RSS feedarxiv.org

Study Finds “Narrative Captivity” in Multi-Turn LLM Advice

Summary

As people use large language models for advice about ethically charged interpersonal conflicts, a new study examines whether narration alone can change a model’s judgment across multiple turns. The authors define “narrative captivity” as a failure mode in which a model treats an unopposed, one-sided account as complete and adopts the narrator’s interpretation without seeking missing perspectives. They create a benchmark containing 5,078 interpersonal-conflict scenarios across six moral dimensions and test 17 LLMs. Across the models, final judgments after multi-turn narration shift by an average of 25 percentage points relative to matched single-turn baselines, indicating that the effect is widespread in the evaluated setting. Stage-level analysis identifies preference optimization as a major contributor. Four inference-time strategies reduce the effect only partially. The authors present the benchmark as a way to encourage AI advisors that retain independent judgment during real-world moral consultations, where one participant’s self-justifying account may unfold over multiple turns and leave the model with asymmetric information.