Back to News
RSS feedarxiv.org

Regularized Emphatic TD Learning Improves Stability Analysis Under Constant Stepsizes

Summary

The paper examines whether emphatic temporal-difference learning (ETD) remains stable when off-policy updates use constant stepsizes. In an ergodic two-state counterexample, the ETD mean map contracts, yet the sampled update product has a positive top Lyapunov exponent, showing that mean stability does not determine sampled constant-stepsize behavior. Regenerative-cycle analysis separates this instability signal from the infinite variance of the follow-on trace. The authors introduce regularized emphatic TD (RETD), which leaves the trace and importance ratios unchanged, stores the emphatic TD signal in a leaky scalar state, and applies a delayed correction. RETD's raw equilibrium is shifted affinely from ETD's, but single- and two-regularization readouts recover the ETD fixed point exactly. The paper proves almost-sure convergence for harmonic diminishing stepsizes and gives a conditional constant-stepsize moment-contraction result based on a Markovian random-product bound. RETD has certified negative exponents on the two-state example and one Baird point, while the positive Baird ETD sign remains numerical. Paired experiments with 10,000 runs support the claimed separation, fixed-point recovery, a nonmonotone stability region, and task dependence. The method changes post-shock dynamics but does not reduce the shared follow-on-trace variance.

Regularized Emphatic TD Learning and Stability | Benpay.ai Board