MEGA Hub

Primitive-Driven Compositional Forensic Visual Prompting for Open-World Face Anti-Spoofing

Authors

Do you know Fangling Jiang?You can claim authorship or link another user.Do you know Qi Li?You can claim authorship or link another user.Do you know Bing Liu?You can claim authorship or link another user.Do you know Weining Wang?You can claim authorship or link another user.Do you know Quilin Huang?You can claim authorship or link another user.Do you know Zhenan Sun?You can claim authorship or link another user.Do you know Ming-Hsuan Yang?You can claim authorship or link another user.

Abstract

Open-world face anti-spoofing must address both covariate and semantic shifts: source and target domains differ in imaging conditions, while target domains contain diverse attack types absent from training. Existing prompt-based approaches often express spoofing through category semantics or language guidance, which is effective for modeling high-level concepts but is less suited to explicitly capturing the evolving fine-grained and spatially heterogeneous forensic evidence of unseen attacks. Motivated by the hypothesis that many unseen attacks can be characterized by new combinations of recurring visual cues, we propose a compositional forensic visual prompt learning framework that operates entirely in the visual feature space. Built on a frozen ViT-based vision foundation model, the framework employs patch-aware attention to refine a shared set of learnable micro-forensic primitives into localized forensic evidence units derived from image patches. Class-specific global contextual prompts then provide input-dependent routing weights that adaptively select and compose these primitives into compositional forensic visual prompts for real/spoof discrimination. The primitives are not assigned predefined semantic meanings; instead, their specialization and reuse emerge from shared parameterization and joint optimization across categories. Extensive experiments on nine open-world protocols demonstrate state-of-the-art performance, strong cross-domain generalization, and robust adaptation to unseen attacks.

Community

00