MEGA Hub

When Diffusion Models Forget Who You Are: Identity Preservation in Face Inpainting under Large Occlusions

Authors

Do you know Feng Ding?You can claim authorship or link another user.Do you know Shuhuai Xie?You can claim authorship or link another user.Do you know Yue Zhou?You can claim authorship or link another user.Do you know Yulan Zhang?You can claim authorship or link another user.Do you know Guopu Zhu?You can claim authorship or link another user.Do you know Mengyao Xiao?You can claim authorship or link another user.

Abstract

Face inpainting with diffusion models has recently achieved impressive visual quality, yet preserving identity fidelity under significant occlusion and conflicting text guidance remains a major challenge. To address this issue, we present Reference Semantic Inpainting for Face (ReSem-Face), a cascaded diffusion framework that introduces an explicit identity-conditioned semantic prior for multi-reference face inpainting. Our approach distills representative identity features from multiple references to reconstruct missing semantic regions, which then guide the diffusion process through a multi-stream conditioning architecture. This design provides strong semantic constraints when pixels are absent and stabilizes identity reconstruction while remaining compatible with prompt-driven edits. Experiments on CelebAHQ-IDI-5 and VGGFace2 demonstrate that ReSem-Face yields more reliable identity-preserving completion under severe semantic masks and improves text-controlled editing quality compared with representative baselines.

Community

00