MEGA Hub

Diffusion Models are Open-World Affordance Learners: Leveraging Generative Priors for 3D Affordance Learning

Authors

Do you know Hanqing Wang?You can claim authorship or link another user.Do you know Zhenhao Zhang?You can claim authorship or link another user.Do you know Kaiyang Ji?You can claim authorship or link another user.Do you know Mingyu Liu?You can claim authorship or link another user.Do you know Wenti Yin?You can claim authorship or link another user.Do you know yuchao chen?You can claim authorship or link another user.Do you know Zhirui Liu?You can claim authorship or link another user.Do you know Xiangyu Zeng?You can claim authorship or link another user.Do you know Tianxiang Gui?You can claim authorship or link another user.Do you know Hangxing Zhang?You can claim authorship or link another user.Do you know Jiahao Yuan?You can claim authorship or link another user.Do you know Zhiqing Cui?You can claim authorship or link another user.Do you know Jiaxin Liu?You can claim authorship or link another user.Do you know Zhiyuan Ma?You can claim authorship or link another user.Do you know Hui Xiong?You can claim authorship or link another user.

Abstract

3D affordance grounding aims to understand how diverse objects can be manipulated, making it a cornerstone of embodied interaction. However, prior works struggle to generalize to out-of-distribution, open-world scenarios, leaving a critical gap between limited dataset performance and real-world application needs. Inspired by the saying: \textit{\textbf{``What I can not create, I do not understand''}}, we find generative models can generate semantically valid HOI images, which indicates inherent encoding of affordance concepts. Building on this insight, we propose DAG, the first innovative diffusion-based 3D affordance grounding framework that extracts general affordance knowledge from text-to-image diffusion models for 3D affordance prediction. Specifically, we extract the affordance priors from a diffusion model to encode HOI priors, and design an affordance block with a multi-source affordance decoder for dense 3D affordance prediction. Extensive experiments show that DAG consistently outperforms state-of-the-art methods and exhibits strong open-world generalization, even in the challenging one-shot setting. The code of our method is released on \textcolor{blue}{\textit{https://github.com/hq-King/DAG}}.

Community

00