MEGA Hub

Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

Authors

Do you know Sreyan Ghosh?You can claim authorship or link another user.Do you know Arushi Goel?You can claim authorship or link another user.Do you know Kaousheik Jayakumar?You can claim authorship or link another user.Do you know Lasha Koroshinadze?You can claim authorship or link another user.Do you know Nishit Anand?You can claim authorship or link another user.Do you know Siddharth Gururani?You can claim authorship or link another user.Do you know Hanrong Ye?You can claim authorship or link another user.Do you know Pritam Biswas?You can claim authorship or link another user.Do you know Yuanhang Su?You can claim authorship or link another user.Do you know Ehsan Hosseini-Asl?You can claim authorship or link another user.Do you know Sang-gil Lee?You can claim authorship or link another user.Do you know Zhifeng Kong?You can claim authorship or link another user.Do you know Jaehyeon Kim?You can claim authorship or link another user.Do you know Sungwon Kim?You can claim authorship or link another user.Do you know S Sakshi?You can claim authorship or link another user.Do you know Ramani Duraiswami?You can claim authorship or link another user.Do you know Dinesh Manocha?You can claim authorship or link another user.Do you know Andrew Tao?You can claim authorship or link another user.Do you know Mohammad Shoeybi?You can claim authorship or link another user.Do you know Bryan Catanzaro?You can claim authorship or link another user.Do you know Ming-Yu Liu?You can claim authorship or link another user.Do you know Wei Ping?You can claim authorship or link another user.

Abstract

We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over long and complex real-world (audio-visual) videos. To support this, we make three key contributions: (i) Audio-Visual-Skills, a large-scale collection of real-world videos with ~7M caption and question-answer training instances designed to emphasize temporal, compositional, and cross-modal audio-visual reasoning; (ii) a novel three-stage curriculum that progressively trains the model from short-range perception to long-horizon multi-event reasoning; and (iii) Temporal Audio-Visual Interleaved Chain-of-Thought, a reasoning framework that explicitly grounds intermediate reasoning steps to timestamps in long audio-visual streams, improving temporal alignment and interpretability. Extensive experiments across 15+ audio-visual, omni-modal, audio, and vision benchmarks show that AV-Flamingo outperforms similarly sized open models by clear margins and remains highly competitive with, and in some cases surpasses, much larger open-weight and closed models, particularly on long and complex real-world audio-visual understanding and reasoning tasks. Beyond benchmark performance, AV-Flamingo exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability.

Community

00