The paper presents TRACER, a multi-turn user simulator designed to reproduce how user intent evolves and how interactions unfold, rather than merely generating plausible individual replies. The system is trained in two stages: supervised fine-tuning on real user dialogues, followed by multi-turn reinforcement learning. Its reinforcement-learning stage combines hierarchical rewards for outcomes and trajectories with deviation-aware advantage modulation, addressing sparse rewards and credit assignment in long conversations. On real customer-service sessions divided into reference cohorts, TRACER-7B reportedly improves conversion F1 by 11.4 points over the strongest baseline. It also records the lowest group-level conversion-rate error and semantic trajectory distance among the compared systems, and generalizes to out-of-distribution scenarios. In human Turing tests, people identify generated conversations with accuracy close to chance, supporting the conversations' perceived naturalness. The authors then introduce the Dynamic Marketing Benchmark, which evaluates language models through simulated interactions using both persuasion effectiveness and response quality. Results from this benchmark indicate that better response quality does not necessarily produce higher conversion rates. The work therefore frames behavioral consistency and interaction outcomes as distinct evaluation targets for interactive AI, while the reported results remain those of the paper's experiments.
AI News
The latest AI releases, research, products, and industry updates.
Loading...