MEGA Hub

Diverse-Intent Multi-Turn Fashion Image Retrieval

Authors

Do you know Mingqiang Tang?You can claim authorship or link another user.Do you know Haokun Wen?You can claim authorship or link another user.Do you know Meng Liu?You can claim authorship or link another user.Do you know Yupeng Hu?You can claim authorship or link another user.Do you know Weili Guan?You can claim authorship or link another user.Do you know Xuemeng Song?You can claim authorship or link another user.

Abstract

Real-world fashion search involves interactive retrieval across multiple turns. However, existing multi-turn retrieval methods are built on a restrictive assumption that every interaction follows the same attribute-editing paradigm, leaving heterogeneous intent transitions unexplored. Moreover, existing approaches often rely on textification to bridge multimodal queries and visual retrieval, which may lose fine-grained visual cues. To address these gaps, we introduce DIM-Fashion, a benchmark of 26K multi-turn sessions constructed from 13 fashion retrieval datasets across 7 tasks, featuring diverse intent transitions and rollback behaviors. We further propose FashionAM, an MLLM-VLP framework that directly aligns multimodal conversational queries with a fashion-oriented gallery embedding space, avoiding intermediate textification. Extensive experiments demonstrate the effectiveness of FashionAM over existing approaches. The dataset and code will be made publicly available upon acceptance.

Community

00