MEGA Hub

UniVerse: Benchmarking and Enhancing LALMs on Culturally Inclusive Low-Resource Music Understanding

Authors

Do you know Ziya Zhou?You can claim authorship or link another user.Do you know Shangda Wu?You can claim authorship or link another user.Do you know Shenyang Xu?You can claim authorship or link another user.Do you know Yutong Zheng?You can claim authorship or link another user.Do you know Dafang Liang?You can claim authorship or link another user.Do you know Suin Chung?You can claim authorship or link another user.Do you know Danbinaerin Han?You can claim authorship or link another user.Do you know Junyan Jiang?You can claim authorship or link another user.Do you know Yongyi Zang?You can claim authorship or link another user.Do you know Ruibin Yuan?You can claim authorship or link another user.Do you know Rongxiu Zhong?You can claim authorship or link another user.Do you know Shilei Zhang?You can claim authorship or link another user.Do you know Junlan Feng?You can claim authorship or link another user.Do you know Jinglei Liu?You can claim authorship or link another user.Do you know Haotian Zhou?You can claim authorship or link another user.Do you know Zijin Li?You can claim authorship or link another user.Do you know Dasaem Jeong?You can claim authorship or link another user.Do you know Wei Xue?You can claim authorship or link another user.Do you know Yike Guo?You can claim authorship or link another user.

Abstract

Recent advances in large audio-language models (LALMs) have significantly improved performance in tasks such as music captioning, genre classification, and sound event detection. However, limited attention has been paid to improving their adaptability across diverse musical traditions, particularly folk music rooted in distinct cultural contexts. Folk-music traditions are typically resource-scarce, unevenly represented across regions, and poorly documented. Even when such samples appear in large-scale pre-training, LALMs often fail to capture their structural and stylistic characteristics, partly due to the absence of dedicated evaluation protocols and training solutions. To address these limitations, we introduce UniVerse, a reproducible solution for low-resource music understanding. Specifically, we propose UniVerseBench, a benchmark of 5,042 Q&A pairs across more than 38 cultural and linguistic entities, constructed via an expert-guided yet highly automated pipeline. In parallel, we construct a fully automated, model-generated multi-turn dialogue training dataset UniVerseSet. By training LALMs on UniVerseSet, we systematically adapt and investigate representative multimodal imbalance learning strategies across both dense and Mixture-of-Experts (MoE) architectures. Experimental results indicate that fully automated data curation combined with imbalance-aware training yields non-trivial improvements, but models still struggle to capture fine-grained acoustic features, indicating a gap between surface-level alignment and deep musical comprehension.

Community

00

Publication notes

Author note
21 pages, 7 figures, 8 tables