MEGA Hub

DAVSS: Distilled Audio-Visual State Space Models

Authors

Do you know Saurabhchand Bhati?You can claim authorship or link another user.Do you know Mrudula Athi?You can claim authorship or link another user.Do you know Amit S. Chhetri?You can claim authorship or link another user.Do you know James Glass?You can claim authorship or link another user.

Abstract

State-space models (SSMs) distilled from transformer teachers combine the performance of transformers with the efficiency of SSMs. We extend the Transformer-SSM knowledge distillation to a multimodal setting and propose the Distilled Audio-visual State-Space (DAVSS) model. The DAVSS model, 14M parameters, is 12 times smaller compared to transformer-based models such as CAV-MAE, and still outperforms them. DAVSS improves over the existing audio-visual models by: 1) Finer input resolution: using smaller patch sizes process the input, compensating for the smaller model size by increasing input sequence lengths. This is supported by the observation that a larger patch size results in lower performance. 2) Deeper joint modeling: utilizing a larger portion of the model (30%) for joint audio-visual processing, compared to <5% in CAV-MAE, enabling deeper cross-modal interaction without significantly increasing the computational cost associated with the concatenated audio-visual tokens.

Community

00