MEGA Hub

DreamHand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery

Authors

Do you know Yufei Liu?You can claim authorship or link another user.Do you know Xixi Wang?You can claim authorship or link another user.Do you know Hao Li?You can claim authorship or link another user.Do you know Ganlong Zhao?You can claim authorship or link another user.Do you know Kaitong Cai?You can claim authorship or link another user.Do you know Chengkai Jin?You can claim authorship or link another user.Do you know Chunxiao Liu?You can claim authorship or link another user.Do you know Jianbo Liu?You can claim authorship or link another user.Do you know Siyuan Huang?You can claim authorship or link another user.Do you know Xingang Pan?You can claim authorship or link another user.Do you know Hongsheng Li?You can claim authorship or link another user.

Abstract

Egocentric video offers scalable manipulation data for embodied AI, yet recovering metric 3D hand trajectories remains challenging due to severe object occlusion and frequent out-of-sight gaps. Existing single-frame and windowed temporal regressors fail when hand shortly leaves the frame, while recent video diffusion models (VDMs) rely on heavy, stochastic multi-step sampling as pixel-space renderers. We instead repurpose VDM into a deterministic geometry encoder. A single forward pass over the clean latent exposes scene content beyond current observations, including occluded and out-of-sight hands. We introduce DreamHand, an offline clip-level framework that extracts features via a Deterministic Clean-Latent Encoder and decodes them with a Bidirectional Spatiotemporal Decoder. DreamHand recovers continuous bimanual trajectories with metric placement and no external detector, while a Ray-Based Camera Solver supports a second configuration that needs no test-time camera intrinsics. Across five egocentric benchmarks, DreamHand sets a new state of the art, cutting MPJPE-p by 30% on occlusion-heavy ARCTIC and 40% on HOT3D. These gains reach 46%-61% once out-of-sight hands are included in the evaluation, offering a scalable path from everyday human video to robot manipulation data.

Community

00

Publication notes

Author note
Project Page: https://ggxxii.github.io/dreamhand/