MEGA Hub

Grounding Free-Form Instructions for Fashion Complementary Image Generation

Authors

Do you know Matteo Attimonelli?You can claim authorship or link another user.Do you know Claudio Pomo?You can claim authorship or link another user.Do you know Alessandro De Bellis?You can claim authorship or link another user.Do you know Danilo Danese?You can claim authorship or link another user.Do you know Dietmar Jannach?You can claim authorship or link another user.Do you know Tommaso Di Noia?You can claim authorship or link another user.

Abstract

Fashion complementary image generation (CIG) aims to create garments that stylistically match a seed item based on user intent, making it a natural multimodal grounding problem where models must interpret language in visual context. Existing CIG benchmarks rely on rigid template prompts (e.g., "a photo of a skirt"), failing to reflect natural user queries and obscuring model behavior across levels of linguistic specificity. We introduce fashion complementary image generation with free-form instructions, a multimodal language-grounding setting where a model generates a compatible garment from a seed image and a natural-language instruction. To this end, we enrich three CIG benchmarks with low-, medium-, and high-specificity instructions generated by a vision-language model and validated by human annotators. We instantiate the task with StyleFlow, a Rectified Flow Matching model that jointly conditions on the seed image and instruction within a single multimodal transformer. Across image quality metrics, catalog-alignment analysis, ablations, and human evaluation, StyleFlow consistently produces instruction-aligned and stylistically coherent garments while reducing architectural complexity and inference cost relative to auxiliary-module approaches.

Community

00