MEGA Hub

Learning to Control Coupled-Dynamics Environments with Joint Markov Decision Processes

Authors

Do you know Ege C. Kaya?You can claim authorship or link another user.Do you know Aliasghar Pourghani?You can claim authorship or link another user.Do you know Mahsa Ghasemi?You can claim authorship or link another user.Do you know Vijay Gupta?You can claim authorship or link another user.Do you know Abolfazl Hashemi?You can claim authorship or link another user.

Abstract

Coupled-dynamics environments expose the one-step outcomes that would follow from several possible counterfactual actions under a common realization of exogenous randomness. The ordinary Markov decision process formalism allows one to reason about the marginal law of each action but discards dependence across these counterfactual outcomes. The Joint Markov decision process (JMDP) formalism preserves that dependence. Prior work established the formalism and solved the fixed-policy joint moment evaluation problem in JMDPs. This paper develops optimal-control methods. We define a nonparametric distributional Bellman optimality operator for JMDPs, and prove that when the induced marginal MDP has a unique optimal policy, its iterates converge in Wasserstein distance to the optimal joint return law. For the first two moments, we establish convergence under a weaker condition that permits several mean-optimal actions as long as their tie resolutions share a second-moment fixed point. We also derive sampled targets for neural approximation.

Community

00

Publication notes

Author note
14 pages, 5 figures