Beyond Mode Collapse: Generating Diverse Synthetic Expert Conversations with Generative Flow Networks
Summary
High-quality synthetic data is important for post-training large language models that need to represent varied expert strategies and decisions. The paper argues that directly prompting language models or conditioning them on end-use scenarios often produces low-diversity data concentrated around dominant modes. It proposes training Generative Flow Networks (GFlowNets) to generate latent conversation structures, using a Gaussian mixture density over interaction features such as confusion dynamics and the balance of scaffolding directives. This lets the generator sample expert strategies in proportion to their prevalence in training data. Across tutoring and emotional-support dialogues, the approach provides a better balance of fidelity, mode coverage, and authenticity than reinforcement-learning and end-to-end LLM baselines, while avoiding copying training examples. On three downstream outcome-prediction tasks, classifiers trained on the GFlowNet-generated conversations receive a stronger training signal than classifiers trained with competing synthesis methods.