AI Agents May Be Farther From Recursive Self-Improvement Than Expected
Summary
A Princeton-led, multi-institution team tested whether AI agents could conduct open-ended AI research rather than merely solve engineering tasks with checkable answers. Using a new “shadow evaluation,” the researchers gave Anthropic’s Claude Opus 4.8, running through the open-source OpenClaw software, unpublished research questions drawn from two NeurIPS 2026 submissions. Each agent had six days, $3,000 in Anthropic API credits, GPU resources, virtual computers, and web access to produce a conference-quality paper. The agents reviewed literature, ran hundreds of experiments, and completed much of the required engineering, but the original paper authors rejected both submissions. The agents used weak experimental designs, including tiny synthetic datasets, abandoned promising hypotheses too quickly, failed to rethink approaches after setbacks, and responded to feedback mainly by narrowing claims and adding caveats. They also struggled to allocate time, compute, and tokens and did not consistently follow process instructions. The orchestrator agent did catch occasional hallucinations or misrepresented results from helper agents, and the study found no clear evidence of reward hacking. The researchers argue that reinforcement learning is easier to apply to tasks with automatically verifiable outcomes than to open-ended research requiring taste and judgment. The study is limited by its two-paper sample, the authors’ awareness that the submissions were AI-generated, and researchers’ discretion in designing the evaluation. Its results nevertheless challenge optimistic timelines for recursive self-improvement. The article notes that Anthropic and OpenAI are pursuing systems that can accelerate AI development, while leaving unresolved whether transformative progress can arise from steadily improving narrow, measurable tasks without creative research breakthroughs.