MEGA Hub

Schrödinger's Cat: Probabilistic Representation and Prediction of Potential Scene Kinematics

Authors

Do you know Timy Phan?You can claim authorship or link another user.Do you know Jannik Wiese?You can claim authorship or link another user.Do you know Björn Ommer?You can claim authorship or link another user.

Abstract

Predicting how a scene may evolve from partial observations requires reasoning about multiple possible futures rather than committing to a single trajectory. Existing approaches either generate appearance-dominated video predictions or sample a small number of trajectories without explicitly modeling the distribution of possible motion. We introduce Goal-Aware Representations of Future kInEmatic Latent Distributions (GARFIELD), a probabilistic model of scene kinematics that learns a structured spatio-temporal latent representation of the distribution over possible futures given an image and optional spatio-temporally sparse constraints. The same latent representation enables both joint sampling of all trajectories and direct access to the underlying motion distribution through an efficient deterministic density decoder. As a result, uncertainty about future motion can be localized to specific scene elements and timesteps and progressively refined through additional constraints. Experiments demonstrate strong motion planning performance competitive with large video generation models while sampling trajectories $97\times$ faster. Our method further estimates motion densities two orders of magnitude faster than Monte-Carlo sampling from motion generation models, enabling interactive exploration and uncertainty-aware planning.

Community

00

Publication notes

Author note
Accepted at ECCV 2026. Project page: https://compvis.github.io/schroedingers_cat