Back to News
RSS feedarxiv.org

AI Agents Are Vulnerable to Radicalization

Summary

This study examines whether large language models can manipulate one another’s beliefs through simulated conversations between two agents. A target LLM role-plays a human persona defined by demographic and psychological attributes, while an influencer LLM tries to make the target’s beliefs more extreme. The researchers test two pathways: resonance, which reinforces a belief the target already holds, and persuasion, which promotes a belief the target initially considers unimportant. Across affective and behavioral measures, both pathways produce radicalization, but resonance has consistently stronger effects. The impact of specific tactics, including sycophancy and unverified claims, varies by metric rather than showing one stable pattern. The study also finds that resonance can spread to related beliefs, pointing to interconnected belief structures within AI agents. The findings raise concerns about personalized agents and multi-agent systems that may be especially vulnerable when messages align with their existing beliefs.