AlphaDiverse Trains Local Agents for Diverse Alpha Factor Mining
Summary
AlphaDiverse is a framework for automating alpha factor mining with local large language model agents. Existing multi-agent systems often depend on external APIs, which limits control over cost, availability, and confidentiality, while long research loops can repeatedly favor a small number of successful economic mechanisms. The framework addresses these problems by combining a multi-agent alpha research system, diverse research-path collection, and post-training for local Planner and Realizer agents. It generates complementary plan portfolios and varies research environments across loops to collect broader traces. Those traces are used for supervised fine-tuning, followed by a joint Group Relative Policy Optimization procedure that optimizes predictive quality and diversity of contributions. Research feedback is restricted to inner-period data, and a frozen final model is evaluated on later outer-period data to avoid test-set tuning. Experiments on four Chinese stock universes report competitive prediction together with broader exploration.