MEGA Hub

Long-Term Sequential Decision Making under Risk

Authors

Do you know Irmaan?You can claim authorship or link another user.Do you know Mirzanejad?You can claim authorship or link another user.Do you know Nadjet Bourdache?You can claim authorship or link another user.Do you know Abdel-Illah Mouaddib?You can claim authorship or link another user.

Abstract

We study finite-horizon MDP planning under \emph{root-based} (resolute) risk objectives that apply a rank-dependent functional to the distribution of total returns. Such objectives are non-linear in the return distribution and generally break Bellman optimality, so direct optimization by scenario-tree enumeration is intractable. We propose \textbf{ERQDP}, an enumeration-free and sampling-free method that solves a rank--quantile surrogate via exact DP (Dynamic Programming), evaluates candidate policies exactly by DP over return Probability Mass Functions (PMFs) on a discretized return grid (with an explicit rounding bound), and refines the surrogate in an anytime loop that reports an explicit upper--lower gap (certificate) for the target objective up to discretization budgets. Across tested benchmarks, ERQDP returns certified solutions or explicit residual gaps, enables fast risk-parameter sweeps with substantial runtime gains, and supports both risk-averse and risk-seeking behaviors.

Community

00

Publication notes

Author note
Accpeted in Forty-Second Annual Conference on Uncertainty in Artificial Intelligence (UAI 2026)