MEGA Hub

ScaleVid: Geometry-Aware Video Object Scaling with Mesh-Free Inference

Authors

Do you know Youze Huang?You can claim authorship or link another user.Do you know Penghui Ruan?You can claim authorship or link another user.Do you know Bojia Zi?You can claim authorship or link another user.Do you know Xianbiao Qi?You can claim authorship or link another user.Do you know Shihao Zhao?You can claim authorship or link another user.Do you know Rong Xiao?You can claim authorship or link another user.

Abstract

Geometry-aware video object scaling aims to anisotropically resize the object along object-centric axes while preserving geometric plausibility, temporal coherence, and background consistency. Existing text-guided methods mainly operate in the 2D image plane, while depth-guided approaches provide coarse control and mesh-based methods require costly 3D reconstruction. We present a progressive two-stage training framework that decouples geometry-aware foreground transformation from background preservation and realistic video composition, without mesh-pixel alignment and explicit 3D reconstruction at inference. In both stages, geometrically perturbed pseudo-sources are constructed from real videos, while the original complete videos are retained as reconstruction targets. The first stage uses planar transformations to learn robust foreground-background composition, whereas the second introduces object-centric 3D deformation guidance for geometry-aware scaling. This pseudo-source reconstruction formulation enables real-video synthesis without paired real-world scaling targets. We construct complementary paired-geometry and real-background benchmarks and further evaluate on in-the-wild videos. Extensive experiments demonstrate superior geometric consistency, foreground fidelity, and background preservation, together with faster and more practical inference than methods requiring explicit 3D reconstruction.

Community

00