MEGA Hub

Vera: Identity-Faithful Human Subject-to-Video Generation

Authors

Do you know Yulong Xu?You can claim authorship or link another user.Do you know Xinyue Liu?You can claim authorship or link another user.Do you know Shujuan Li?You can claim authorship or link another user.Do you know huafeng shi?You can claim authorship or link another user.Do you know Yan Zhou?You can claim authorship or link another user.Do you know Jiwen Liu?You can claim authorship or link another user.Do you know Xintao Wang?You can claim authorship or link another user.Do you know Yu Shen Liu?You can claim authorship or link another user.Do you know Huaibo Huang?You can claim authorship or link another user.

Abstract

Subject-to-video (S2V) generation has made substantial progress in preserving reference subjects across diverse categories, yet generic subject consistency remains insufficient for human-centric generation. A video may appear globally consistent while identity-critical human details still drift across frames, poses, and interactions. This issue becomes more severe in multi-person scenarios, where incorrect identity-role binding leads to subject confusion, attribute swapping, and excessive copying of reference-specific appearance cues. We propose Vera, a unified human-centric S2V framework for single- and multi-person generation. We first construct a million-pair identity-aligned human image-video dataset through person-level cross-clip retrieval, providing explicit identity correspondence and diverse references. Built on this dataset, Vera introduces two complementary designs. Identity-Focal Masked Supervision (IFMS) strengthens identity-aware learning with spatially focused supervision while reducing interference from irrelevant artifacts. Reference-Aware Layer-wise Attention (RALA) regulates how video tokens interact with reference identity cues in the DiT backbone, preserving stable identity anchors and enhancing layer-aware identity readout. Extensive experiments demonstrate that Vera improves human identity consistency, multi-person subject binding, and motion naturalness, while reducing identity confusion and excessive reference-image copying.

Community

00