AdvRole is an adversarial context-rewriting framework for reinforcement learning with large-language-model role-playing agents. The paper identifies a distributional bottleneck in conventional methods: they train on a fixed scenario pool collected before learning, even though the agent’s weaknesses change as it improves. AdvRole alternates between an Actor that learns to role-play and a Rewriter that edits character profiles and dialogue contexts into hard, actor-specific scenarios. The Rewriter receives a performance-gap reward that favors rewrites causing the current Actor to score lower than it does on the original scenario. This makes the scenario pool evolve with the Actor and continually target under-mastered regions of the character-context space. Experiments across three role-playing benchmarks covering English and Chinese, plus a new multilingual benchmark released by the authors, show that AdvRole consistently outperforms the evaluated baselines.
AI News
The latest AI releases, research, products, and industry updates.
Loading...