Training Idea-Level Critics to Make ML-Evolving Agents More Verification-Efficient
Summary
As large language models enable self-evolving agents to work on AI for machine learning, empirical verification remains a bottleneck because each training and evaluation cycle can be computationally expensive. The paper introduces specialized idea-level critic models that predict whether a proposed machine-learning modification will improve the current solution, allowing agents to spend verification resources on more promising ideas. The critics are trained with supervised fine-tuning on high-quality critiques synthesized by Gemini-3.1-Pro, followed by GRPO to improve predictive accuracy. In static idea evaluation, the trained critics outperform Gemini-3.1-Pro, and the advantage carries over to agent inference, continual learning, and policy training. During inference-time evolution, they improve final solution quality under the same verification budget by selecting better candidates, with continual learning providing additional gains. During policy training, the critics act as learned reward models: empirical verification is reserved for uncertain cases, enabling substantially more policy updates with the same resources. The results support idea-level criticism as a way to help ML agents discover stronger solutions and proposal policies when verification is limited.