Back to News
RSS feedarxiv.org

Self-Discovering Reinforcement Learning: Is Learning History an Asset or a Burden?

Summary

The paper examines whether learning history enables or constrains reinforcement-learning systems intended for recursive self-improvement. It presents what it describes as the first causal mechanistic audit of Disco103, a self-discovered RL update rule previously reported to outperform PPO on benchmark performance. The audit is organized around five conditions associated with the “Era of Experience”: extended horizons, grounded reward scales, continuing data streams, within-lifetime change, and exploration depth. The researchers pin, freeze, and transplant recurrent states while keeping meta-parameters fixed, allowing them to separate the content of history from the mechanism that maintains it. They report that recurrent history expands the usable reward range to six orders of magnitude, compared with three orders under zero-pinning. They also find that mismatched history is costly largely because the imported state is continually clamped; allowing it to evolve naturally reduces the burden. In changing environments, controlling replay retention removes the apparent adaptation advantage over DQN, indicating that external data turnover can confound conclusions about internal plasticity. The findings are checked using capability thresholds and transferred to a second rule, OPEN. The authors present the results as a basis for auditing learning dynamics in future self-evolving RL algorithms.