MEGA Hub

P3: Probabilistic Policy Propagation for Stable VAE-Based Robot Learning

Authors

Do you know Liyun Yan?You can claim authorship or link another user.Do you know Jianming Ma?You can claim authorship or link another user.Do you know Yang Zhang?You can claim authorship or link another user.Do you know Shengcheng Fu?You can claim authorship or link another user.Do you know Zhanxiang Cao?You can claim authorship or link another user.Do you know Keqi Zhu?You can claim authorship or link another user.Do you know Yizhi Chen?You can claim authorship or link another user.Do you know Yue Gao?You can claim authorship or link another user.

Abstract

Variational Autoencoders are widely used to encode high-dimensional and noisy observations in robotics. However, their stochastic latent creates a mismatch with Proximal Policy Optimization (PPO): an effective policy marginalizes over the latent distribution, whereas former implementations estimate its probability ratio and KL divergence using only one latent sample. We identify a fundamental but overlooked theoretical cause: naive single-sample approximations in stochastic latent space induce significant variance and bias in the surrogate loss. To address this, we introduce P^3 (Probabilistic Policy Propagation), a distribution-aware optimization framework for VAE-based policies. $P^3$ couples moment-based probabilistic method for stable and efficient learning with sampling-based calibration for robust policy behavior under latent uncertainty. In our experiments, P^3 boosts data efficiency from 64.6% to >96%, reduces convergence steps by >20%. Furthermore, P^3 is evaluated on challenging humanoid parkour tasks and shows an effective foundation for VAE-based PPO. Code is available at https://github.com/ylyem9x/P3_Open.

Community

00