Back to News
RSS feedarxiv.org

Failure Happens Before Drift: Values in LLM Agent Societies

Summary

The study presents a World Values Survey-grounded framework for testing whether culturally diverse LLM agents can faithfully represent conflicting human value systems over time. Agents with different communication styles held longitudinal discussions across about 4,000 conversations involving 1,200 personas, 15 topics, and three models: GPT-4o, Gemini-2.5-Flash, and Gemma-4-E4B. The researchers measured value faithfulness, value drift, and conversational realism. More than 50% of personas failed to express their assigned WVS profiles from the beginning, while 2-7% drifted after repeated conversations. Removing demographic details improved faithfulness for some models but did not eliminate systematic differences between simulated and assigned value distributions. Compared with human discussions, the simulated dialogues were often semantically diverse but stylistically repetitive, revealing a different balance between consistency and variation. The authors conclude that current LLM agents can produce plausible conversations, but remain limited proxies for preserving diverse human value profiles over time.