MEGA Hub

Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency

Authors

Do you know Guo An?You can claim authorship or link another user.Do you know Zijing Wu?You can claim authorship or link another user.Do you know Honghua Dong?You can claim authorship or link another user.Do you know Yuhao Yan?You can claim authorship or link another user.Do you know Zixuan Gui?You can claim authorship or link another user.Do you know Haochong Chen?You can claim authorship or link another user.Do you know Shanzhao Ruan?You can claim authorship or link another user.Do you know Xiang Wang?You can claim authorship or link another user.Do you know Yurong Ling?You can claim authorship or link another user.Do you know Qi Tian?You can claim authorship or link another user.

Abstract

Joint-embedding predictive architectures (JEPAs) learn world models that predict in a compact latent space rather than in pixels, reducing the pressure to model nuisance appearance. Yet this provides no guarantee against visual perturbations: they can still alter the encoded representation and affect subsequent action-conditioned predictions. Bisimulation captures this requirement precisely: two observations should be treated as the same state only when their action-conditioned consequences agree. Guided by this criterion, we introduce Action-Conditioned Predictive Consistency (ACPC), a diagnostic that measures how far a clean history and a visually perturbed view of it diverge after being rolled forward under the same action sequence. We prove that this divergence bounds the perturbation-induced change in multi-step prediction error and planner cost. Building on pairwise ACPC, we define two complementary measures: the Invariance Radius (IR) summarizes clean-perturbed rollout spread, while the Separation Rate (SR) checks whether different states remain distinguishable after rollout. Experiments on four visual control tasks show that pairwise ACPC predicts perturbation-induced prediction and cost changes. On LeWM, the IR-SR screen transfers across tasks, and the joint diagnostic remains informative under blur and resize. PLDM exhibits similar diagnostic trends under a different architecture.

Community

00