Back to News
RSS feedarxiv.org

Risk-Averse Online POMDP Planning with CVaR-Based Immediate Costs

Summary

Online POMDP planners that optimize expected cumulative cost can overlook dangerous states within the current belief. Existing risk-averse approaches apply CVaR to the value function, addressing trajectory-level risk but leaving this within-belief immediate-cost risk unresolved and requiring specialized algorithms. The paper instead applies CVaR directly to the state-dependent immediate cost under the belief at each step. It retains the standard expected cumulative return objective, so the resulting problem has an ordinary MDP structure and existing expectation-based POMDP planners can be adapted by changing only cost computation. The authors inherit finite-time guarantees for policy evaluation and sparse sampling, with estimation error independent of the risk level. They also prove a finite-time bound between the particle-belief MDP surrogate and the original POMDP, producing an end-to-end guarantee from the true POMDP value to the algorithmic estimate. When the risk level reaches its risk-neutral limit, the formulation reduces to standard expectation-based planning.