AI Agents Rediscover 62.7% of Findings in Open-Ended Science Tasks
Summary
The preprint investigates whether AI agents can conduct open-ended scientific discovery rather than solve tasks with clearly defined metrics. It evaluates agents in Station, an open-world environment where multiple agents simulate a scientific ecosystem. The authors add a Supervisor mechanism and periodic Meta Reflection to encourage continued exploration when intermediate measures are unavailable. They create tasks from three recent ICLR oral papers, provide only each paper’s main research question, withhold the results, and disable web access. Performance is measured by the share of the original findings, divided into individual criteria, that the agents rediscover. Station recovers 62.7% of the criteria on average, compared with 15.4% for Codex Multiagent-v2 and 14.4% to 20.6% for AI Scientist-v2. Ablation and behavioral analyses indicate that the two mechanisms together improve both research coverage and continuity. In two additional tasks without oracle papers, some agent discoveries closely match findings later reported by researchers after the relevant knowledge cutoff date. The authors interpret these results as evidence that an appropriately designed environment can help agents make meaningful autonomous progress on open-ended research, while the paper remains a preprint and reports results from a defined task set rather than general scientific capability.