MEGA Hub

RAGMesh with FaME-G2E: Long-Form Text-Driven 3D Face Generation and Editing

Authors

Do you know Hao Li?You can claim authorship or link another user.Do you know Ju Dai?You can claim authorship or link another user.Do you know Feng Zhou?You can claim authorship or link another user.Do you know Mengting Shi?You can claim authorship or link another user.Do you know Haofei Wang?You can claim authorship or link another user.Do you know Zhen Song?You can claim authorship or link another user.Do you know Wei Zhou?You can claim authorship or link another user.Do you know Lei Li?You can claim authorship or link another user.Do you know Junjun Pan?You can claim authorship or link another user.

Abstract

Text-driven 3D face generation and editing remains challenging due to the difficulty of translating long-form descriptions into fine-grained facial geometry. Existing methods primarily align global textual semantics with facial structures but often struggle to capture subtle local deformations, such as eyebrow tension, cheek contraction, and asymmetric mouth motions, resulting in limited geometric fidelity and editing precision. To facilitate fine-grained text-driven facial modeling, we first construct FaME-G2E, a large-scale multimodal dataset containing detailed text--mesh annotations and paired text--blendshape samples for unified 3D facial generation and editing. Based on this dataset, we propose RAGMesh, a retrieval-augmented framework that leverages text-correlated geometric priors to improve high-fidelity facial synthesis and editing. Specifically, the Multi-Scale Retrieval Fusion (MSRF) module retrieves semantically consistent global and regional facial priors and fuses them in the blendshape space, suppressing conflicting local deformations while preserving coherent deformation patterns. Furthermore, we introduce Adaptive RAG-guided Supervision (AdaRAGS), a region-aware constraint that explicitly aligns textual semantics with corresponding facial regions, enhancing regional controllability and editing accuracy. Extensive experiments on FaME-G2E demonstrate that RAGMesh achieves superior performance over state-of-the-art methods in local geometric accuracy, text-guided controllability, regional editing precision, and inference efficiency. Video demo is available at https://youtu.be/Yr0_XkpWcNk, and the source code and dataset will be released upon paper acceptance.

Community

00