Back to News
RSS feedwww.dwarkesh.com

AI Researchers Debate the Path to Recursive Self-Improvement

Summary

Dwarkesh Patel interviews John Schulman, Beren Millidge and Charlie O’Neill about whether current AI systems are approaching recursive self-improvement and what could prevent a rapid takeoff. They identify persistent weaknesses in generalization, long-horizon judgment, objective specification, continual learning and the sim-to-real gap as possible bottlenecks. Schulman argues that models may repeatedly appear close to AGI before users encounter weaknesses in self-checking and judgment, while Millidge highlights the difficulty of making systems propose and learn their own objectives. The guests distinguish cumulative AI research tasks, where discoveries can be added to a training pipeline, from non-stationary real-world work that requires adapting to people, institutions and changing information. They expect automated AI researchers to be trained through a mixture of human feedback, multi-step research environments, synthetic data and iterative correction, but disagree about how far these systems can generalize beyond well-specified objectives. The discussion says distillation can weaken the concentration of model providers, although useful distillation depends heavily on realistic prompt distributions and environments rather than benchmark puzzles alone. Deployment data can improve later model generations, but direct continual updates remain difficult because small, repeated updates cause catastrophic forgetting and degrade existing capabilities; large-scale retraining or periodic consolidation is still often needed. The participants also argue that recent RL gains are partly driven by strong mid-training data and high-signal verifiers, with RL changing policies in ways that can produce large behavioral effects despite small parameter updates. RL appears to generalize more reliably across task horizons than across domains, while narrow reward functions can reduce output diversity and encourage reward hacking. On timelines, the guests suggest that capable month-long white-collar agents could emerge within roughly one to three years, while estimates for AI systems dominating computer-based expert work range from about three to ten years, with long-horizon learning and underdeveloped environments remaining major uncertainties.