MEGA Hub

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

Authors

Do you know Haodong Li?You can claim authorship or link another user.Do you know Tianfei Ren?You can claim authorship or link another user.Do you know Xiaoxiao Ma?You can claim authorship or link another user.Do you know Chunmei Qing?You can claim authorship or link another user.Do you know Zhen Fang?You can claim authorship or link another user.Do you know Sipeng He?You can claim authorship or link another user.Do you know Ziyu Guo?You can claim authorship or link another user.Do you know Haoyu Wu?You can claim authorship or link another user.Do you know Juanxi Tian?You can claim authorship or link another user.Do you know Yihang Zou?You can claim authorship or link another user.Do you know Ruichuan An?You can claim authorship or link another user.Do you know Dongzhi Jiang?You can claim authorship or link another user.Do you know Boxue Yang?You can claim authorship or link another user.Do you know Ji Xie?You can claim authorship or link another user.Do you know Xu Huang?You can claim authorship or link another user.Do you know Wenhao Yan?You can claim authorship or link another user.Do you know Jialv Zou?You can claim authorship or link another user.Do you know Zhengrong Yue?You can claim authorship or link another user.Do you know Yaxin Luo?You can claim authorship or link another user.Do you know Xiaotong Li?You can claim authorship or link another user.Do you know Yuzhu Wang?You can claim authorship or link another user.Do you know Junyan Ye?You can claim authorship or link another user.Do you know Jinjing Zhao?You can claim authorship or link another user.Do you know Zehui Chen?You can claim authorship or link another user.Do you know Lin Chen?You can claim authorship or link another user.Do you know Renye Yan?You can claim authorship or link another user.Do you know Feng Zhao?You can claim authorship or link another user.Do you know Pheng-Ann Heng?You can claim authorship or link another user.

Abstract

Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought approaches introduce intermediate plans or visual states, but these representations are typically non-executable or temporally sparse, limiting their ability to instantiate and control the complete spatiotemporal process. To address this limitation, we introduce VideoCoCo, an agentic dual-engine framework in which executable Blender code serves as a process-level chain of thought. Given a text prompt, a coding agent synthesizes a Blender program that explicitly specifies the scene and its temporal evolution. The executable simulation engine runs the program to produce a deterministic spatiotemporal draft, which is subsequently transformed into a photorealistic video by a generative video engine through draft-conditioned editing. This decomposition separates process-level reasoning from high-fidelity visual realization. To adapt the video editor to simulated drafts, we construct VideoCoCo-3K, a curated dataset of draft-instruction-target triplets. VideoCoCo improves the OmniWeaving baseline from 0.475 to 0.558 on PhyGenBench and from 52.18 to 77.88 on VBench-2.0, achieving the best average score on both benchmarks. These results demonstrate that executable code provides an effective, controllable, and inspectable intermediate representation for physically consistent video generation.

Community

00